Skip to content
Source
arXiv Robotics — research abstracts· Ryan Li, Yigit Korkmaz, Erdem B{\i}y{\i}k·· 3 hours agoEditorial score65

Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

Reward-DAgger: Robot-Gated Interactive Imitation Learning with General-Purpose Progress-Based Reward Models

Summary

This paper presents Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. The approach is agnostic to policy architecture, requires no access to policy internals, and achieves better failure-detection accuracy-latency tradeoffs than existing baselines across simulated and real-world tasks.

Source: arXiv Robotics — research abstracts · Read original article ↗

Article text · Original source · English

arXiv:2610.04054v1 Announce Type: new Abstract: Recent advances in robot learning have enabled generalist control policies capable of completing a wide range of tasks. However, their performance degrades when deployed in unseen environments, making it critical to detect failures and teach recovery behaviors. Existing runtime monitoring methods often require task- and policy-specific training or hyperparameter tuning, limiting cross-task deployment and introducing additional overhead during iterative policy updates. We present Reward-DAgger, a robot-gated interactive imitation learning framework that uses dense progress signals from a general-purpose reward model to determine when human intervention is needed. Our approach is agnostic to the underlying policy architecture, requires no access to policy internals, and can be applied across tasks without retuning the gating mechanism. Our results show that Reward-DAgger achieves a better failure-detection accuracy-latency tradeoff than existing runtime monitoring baselines. Across eight simulated and real-world tasks, Reward-DAgger consistently improves the downstream policy's autonomous success rate throughout interactive learning and achieves strong return on human effort, outperforming the baselines in most settings. Importantly, the same gating configuration is used across tasks without task-specific hyperparameter tuning, demonstrating transfer across tasks, environments, and policy architectures. Code and videos are available at https://liralab.usc.edu/reward-dagger.

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).