SUAVE: Unified Video-Action Models via Masked Diffusion
Overview
SUAVE is a unified video-action model using masked diffusion to generate video and actions from language. It functions as a world model with discrete token representation across modalities.
Generated from attributed reports · 50 minutes agoUpdated
Event evidence and corrections
0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.
Reported quantity · actions: 2.5 other · Basis not reported
“ideo and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improve”
Exact source · revision 1Source owner not reported
Reported quantity · time: 1030 other · Basis not reported
“ideo and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improve”
Exact source · revision 1Source owner not reported
Report timeline
Follow attributed reports and material updates.
- arXiv Robotics — research abstractsSUAVE: Unified Video-Action Models via Masked Diffusion
SUAVE is a unified video-action model that uses a masked diffusion transformer to generate video and actions conditioned on language, with all modalities represented as discrete tokens. It can function as a world model, robot policy, or video-action model depending on masked tokens. The model shows strong performance in simulation and real-world tasks, with fast inference and closed-loop control capabilities.
Event attention history
There is not enough continuous observation data to show a trend.
Timezone · UTC
Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).