SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation
SimVLA: Zero-Shot Sim-to-Real VLA Learning for Mobile Manipulation
SimVLA is an end-to-end framework that trains visual-language agents (VLAs) entirely on synthetic simulation data for mobile manipulation. It uses two complementary datasets, SimAction and SimVQA, and is evaluated on real-world tasks, demonstrating zero-shot transfer and outperforming policies trained on real-world demonstrations.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
- Environment
- SimulationOpen source S3
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
s entirely on synthetic simulation data without teleoperation for mobile manipulation. SimVLA is first pre-trained on two complementary simulation-derived datasets: SimAction, a large-scale robot action dataset spanning 35 diverse mobile manipulation tasks, generated by composing atomic skills, and SimVQA, which leverages privileged simulator state to provide spatial, geometric, and subtask-level visual-language supe
Open source S3
Source:arXiv Robotics — research abstracts · arxiv.org