SUAVE: Unified Video-Action Models via Masked Diffusion
SUAVE: Unified Video-Action Models via Masked Diffusion
SUAVE is a unified video-action model that uses a masked diffusion transformer to generate video and actions conditioned on language, with all modalities represented as discrete tokens. It can function as a world model, robot policy, or video-action model depending on masked tokens. The model shows strong performance in simulation and real-world tasks, with fast inference and closed-loop control capabilities.
Source: arXiv Robotics — research abstracts · Read original article ↗
Article text · Original source · English
arXiv:2610.04009v1 Announce Type: new Abstract: Vision-language-action models (VLAs) inherit strong semantic grounding from pretrained vision-language backbones but are typically optimized for predicting actions rather than future observations. They can see and act, but they do not imagine the future before acting. World action models (WAMs) built on video diffusion backbones can imagine but treat language as frozen conditioning on a continuous latent space. Unified models bring these modalities into one architecture, but they either decode autoregressively, one token at a time, or keep video continuous with an auxiliary action head. In this work, we present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss. Simulation and real-world experiments demonstrate two findings. First, a single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift. Together, these results show that masked diffusion is a practical and versatile foundation for unified video-action modeling.
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
Reported numbers
actions
2.5
time
1,030
View original evidence
generates subgoal images and an action chunk spanning one second of motion in 1,030 ms
Open source S5
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. For action-free co-training, the action positions of unlabeled video are filled with mask tokens and excluded from the loss. Simulation and real-world experiments demonstrate two findings. First, a single SUAVE model predicts long-horizon v
Open source S4
ideo and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improve
Open source S5
Source:arXiv Robotics — research abstracts · arxiv.org