Skip to content
arXiv Robotics — research abstracts· Yonghoon Dong, Kyungmin Lee, Changyeon Kim, Jaehyuk Kim, Jinwoo Shin·· 1 days agoEditorial score42

Trust Region Q Adjoint Matching for Stable Off-Policy RL

Trust Region Q Adjoint Matching

Summary

TRQAM improves off-policy reinforcement learning stability by controlling path-space KL divergence via a trust-region parameter. It outperforms prior methods with 68% success in offline RL.

Source: arXiv Robotics — research abstracts · Read original article ↗

Loading article text…

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).