arXiv Robotics — research abstracts· Yonghoon Dong, Minsung Yoon, Jaehyuk Kim, Jungwoo Park, Changyeon Kim, Jinwoo Shin·· 3 days agoEditorial score63
Q-Learning with Scalar Adjoint Matching
Q-Learning with Scalar Adjoint Matching
Summary
This research proposes SQAM, a reinforcement learning method that combines scalar adjoint matching with value penalties to improve fine-tuning of flow policies. By reducing the computational cost of vector-Jacobian products, SQAM achieves performance gains of 18 to 35 percentage points over strong baselines in four hard OGBench domains. The method is also tested on a real bimanual robot, showing improvements over supervised fine-tuning.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
Source:arXiv Robotics — research abstracts · arxiv.org