Skip to content
arXiv Robotics — research abstracts· Yonghoon Dong, Minsung Yoon, Jaehyuk Kim, Jungwoo Park, Changyeon Kim, Jinwoo Shin·· 3 days agoEditorial score63

Q-Learning with Scalar Adjoint Matching

Q-Learning with Scalar Adjoint Matching

Summary

This research proposes SQAM, a reinforcement learning method that combines scalar adjoint matching with value penalties to improve fine-tuning of flow policies. By reducing the computational cost of vector-Jacobian products, SQAM achieves performance gains of 18 to 35 percentage points over strong baselines in four hard OGBench domains. The method is also tested on a real bimanual robot, showing improvements over supervised fine-tuning.

Source: arXiv Robotics — research abstracts · Read original article ↗

Loading article text…

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).