Skip to content
Source
Sergey Levine· @svlevine · X·· 12 days agoEditorial score23

How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors! A fun collaboration with Siemens, led by Brian Zhu, Momen Khalil, Emanuele Poggi from Siemens and @ehharrison4 from Berkeley, with lots of amazing contributors!

Summary

How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors! A fun collaboration with Siemens, led by Brian Zhu, Momen Khalil, Emanuele Poggi from Siemens and @ehharrison4 from Berkeley, with lots of amazing contributors!

Evidence and limits

Published automatically after robotics and source-evidence checks; no manual editorial approval is recorded. Source assertions are not independently verified. Missing information remains not reported.

Environment:
Not reported
Control:
Not reported
Data origin:
Not reported
Source excerpts and review record

No manual editorial approval recorded.

Original source quotation: “How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors!”

Source E1

Original source quotation: “Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning.”

Source E2

Original source quotation: “How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information!”

Source E3
Text

How do we run RL with real-time chunking (RTC)? In this work we figured out how to use a small RL policy with a large robot foundation model, where the RL policy observed more recent images (due to faster inference) and steers the policy toward better behaviors!

A fun collaboration with Siemens, led by Brian Zhu, Momen Khalil, Emanuele Poggi from Siemens and @ehharrison4 from Berkeley, with lots of amazing contributors!

Quoted postE Harrison@ehharrison4
Asynchronous VLA inference reduces inference delay, but breaks the Markovian assumption necessary for RL fine-tuning. How can we enable RL fine-tuning of VLAs with async inference? We introduce ARLI: Asynchronous RL with Intermediate Information! https://async-rl-intermediate-information.github.io/ (1/n)
View the quoted post on X

Source:Sergey Levine · x.com