Skip to content
Trending eventDeveloping

SUAVE: Unified Video-Action Models via Masked Diffusion

1 reports1 reporting sources2 hours agoUpdated

Overview

Event synthesis

SUAVE is a unified video-action model using masked diffusion to generate video and actions from language. It functions as a world model with discrete token representation across modalities.

Generated from attributed reports · 50 minutes agoUpdated

Event evidence and corrections

0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.

Reported quantity · actions: 2.5 other · Basis not reported
Supporting report

“ideo and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improve”

Exact source · revision 1

Source owner not reported

Reported quantity · time: 1030 other · Basis not reported
Supporting report

“ideo and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks. On a real robot, our model generates subgoal images and an action chunk spanning one second of motion in 1,030 ms on an RTX 5090 GPU, sustaining closed-loop control at 2.5 actions per second. Second, pretraining on robot video and co-training on human video substantially improve”

Exact source · revision 1

Source owner not reported

Report timeline

Follow attributed reports and material updates.

10/6
  1. arXiv Robotics — research abstracts
    SUAVE: Unified Video-Action Models via Masked Diffusion

    SUAVE is a unified video-action model that uses a masked diffusion transformer to generate video and actions conditioned on language, with all modalities represented as discrete tokens. It can function as a world model, robot policy, or video-action model depending on masked tokens. The model shows strong performance in simulation and real-world tasks, with fast inference and closed-loop control capabilities.

Event attention history

There is not enough continuous observation data to show a trend.

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).