Skip to content
arXiv Robotics — research abstracts· Yuchen Zhu, Chenyi Xu, Yulin Zhang, Gang Xu, Wentao Zhu·· 1 days agoEditorial score65

Juno: Taming Predictive Latents for Vision-Language-Action Models

Juno: Taming Predictive Latents for Vision-Language-Action Models

Summary

Juno is a unified framework for vision-language-action (VLA) models that addresses three key failures in predictive latent representation: embodiment-specific control mismatch, action learning interference, and teacher miscalibration under distribution shifts. It improves success rates in both simulated and real-world environments, achieving 72.7% on SimplerEnv and retaining 70–75% success under various shifts on real robots.

Source: arXiv Robotics — research abstracts · Read original article ↗

Loading article text…

Source:arXiv Robotics — research abstracts · arxiv.org

Timezone · UTC

Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).