Skip to content
Haozhi Qi· @HaozhiQ · X·· 1 days agoEditorial score75

RPG: Guided Self-Improvement for Embodied Agents

Summary

RPG is a real-to-sim-to-real pipeline that transforms real-world demonstrations into targeted practice for embodied agents. It reconstructs practice tasks from demonstrations, uses demo videos to guide skill improvement, and continuously updates a shared skill system through failure analysis. On real robots, RPG achieves 30/30 successes across three tasks, and in simulation, it reaches 95.0% success across 22 manipulation tasks after 15 rounds of self-improvement.

Source: Haozhi Qi — X Robotics · Read original article ↗

Article text · Original source · English

We have seen quite a few capable agents controlling robots. But an important next question is: how can we build an automatic pipeline for these agents to keep improving themselves?

Check out RPG.

Starting from an offline robot dataset, RPG automatically figures out what capabilities need to be practiced, builds sim environments for those skills (practice tasks, rather than digital twins), and then improves the robot’s skill library through repeated rounds of practice and failure analysis.

Then, the improved system also transfers back to the real world, achieving 30/30 successful hardware trials across three tasks.

Today we use datasets mainly to train robots. In the future, datasets can also tell robots what they should practice next. As robot data becomes increasingly abundant, automatically turning those experiences into targeted practice and continual self-improvement feels is really promising!

Quoted postYen-Jen Wang@wangyenjen
Practice makes perfect. A truly agentic robot shouldn’t just know what to do. It should also figure out how to get better. What if a robot could watch a demonstration, figure out what it needs to practice, and learn how to get better through practice? Introducing RPG: Guided Self-Improvement for Embodied Agents, a real-to-sim-to-real pipeline that turns real-world demonstrations into targeted practice. RPG reconstructs practice tasks from demonstrations, uses demo videos to guide skill improvement, and continually updates a shared skill system from failures. On real robots, RPG achieves 30/30 successes across 3 tasks. In simulation, it reaches 95.0% success across 22 manipulation tasks after 15 rounds of self-improvement. Demo → Reconstruct → Guided Practice → Self-Improve → Go Real
View the quoted post on X

Source:Haozhi Qi · x.com

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).