PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies
This paper presents PLaW-VLA, a method for vision-language-action (VLA) policies that models task-relevant future states in a pretrained prediction-oriented representation space. By focusing on predictive context rather than detailed visual reconstruction, PLaW-VLA achieves improved long-horizon control and generalization on benchmarks like RoboTwin Hard Horizon III and LIBERO-Plus, with lower inference latency compared to generative world-action modeling approaches.
Source: arXiv Robotics — research abstracts · Read original article ↗
Loading article text…
Source:arXiv Robotics — research abstracts · arxiv.org