arXiv Robotics — research abstracts· Yuchen Zhu, Chenyi Xu, Yulin Zhang, Gang Xu, Wentao Zhu·· 1 days agoEditorial score65
Juno: Taming Predictive Latents for Vision-Language-Action Models
Juno: Taming Predictive Latents for Vision-Language-Action Models
Summary
Juno is a unified framework for vision-language-action (VLA) models that addresses three key failures in predictive latent representation: embodiment-specific control mismatch, action learning interference, and teacher miscalibration under distribution shifts. It improves success rates in both simulated and real-world environments, achieving 72.7% on SimplerEnv and retaining 70–75% success under various shifts on real robots.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
Source:arXiv Robotics — research abstracts · arxiv.org