Our Journey Building Vision-Language-Action Robotics
Overview
A robotics team's progress from imitation learning to VLA models, highlighting environmental adaptation challenges.
Generated from attributed reports · 55 minutes agoUpdated
Event evidence and corrections
0 attributed source owners. Ownership does not establish independent confirmation. Quantities are reported separately and are never added together.
Reported quantity · training steps: 100000 other · Basis not reported
“ation Learning Our first baseline used ACT - Action Chunking with Transformers . For the experiment, we trained using: 50 demonstrations 100,000 training steps Batch size: 8 When the cube started near positions well represented in the training demonstrations, the robot could complete the task. But when we changed the cube's initial position, execution became much less reliable. That observation was important. It did”
Exact source · revision 2Source owner not reported
Reported quantity · demonstrations: 50 other · Basis not reported
“ation Learning Our first baseline used ACT - Action Chunking with Transformers . For the experiment, we trained using: 50 demonstrations 100,000 training steps Batch size: 8 When the cube started near positions well represented in the training demonstrations, the robot could complete the task. But when we changed the cube's initial position, execution became much less reliable. That observation was important. It did”
Exact source · revision 2Source owner not reported
Reported quantity · parameters: 450000000 other · Basis not reported
“r First Deployed VLA: SmolVLA Our first major step was SmolVLA . At roughly 450 million parameters , the model was compact enough for us to deploy on an edge device. During our tests, we observed approximately 200 ms response latency from the VLA pipeline. Getting the model running on the robot itself was an important milestone. For the first time, we had the complete flow operating: Camera observation → Language ins”
Exact source · revision 2Source owner not reported
Reported quantity · response latency: 200 other · Basis not reported
“r First Deployed VLA: SmolVLA Our first major step was SmolVLA . At roughly 450 million parameters , the model was compact enough for us to deploy on an edge device. During our tests, we observed approximately 200 ms response latency from the VLA pipeline. Getting the model running on the robot itself was an important milestone. For the first time, we had the complete flow operating: Camera observation → Language ins”
Exact source · revision 2Source owner not reported
Reported quantity · episodes: 100 other · Basis not reported
“rajectories or recordings could affect execution significantly. This taught us another lesson: Robotics AI is not only a model problem. It is also a data problem. Increasing Model Capacity: π0.5 Our next experiment moved to Physical Intelligence's π0.5 . We fine-tuned the approximately 3-billion-parameter model using 100 episodes and evaluated a 5K checkpoint on the physical SO-101 robot. The change was noticeable. C”
Exact source · revision 2Source owner not reported
Reported quantity · checkpoints: 5000 other · Basis not reported
“rajectories or recordings could affect execution significantly. This taught us another lesson: Robotics AI is not only a model problem. It is also a data problem. Increasing Model Capacity: π0.5 Our next experiment moved to Physical Intelligence's π0.5 . We fine-tuned the approximately 3-billion-parameter model using 100 episodes and evaluated a 5K checkpoint on the physical SO-101 robot. The change was noticeable. C”
Exact source · revision 2Source owner not reported
Report timeline
Follow attributed reports and material updates.
- Robotics — Paper and dataset web discoverySignalOur Journey Building Vision-Language-Action Robotics
This article details the progression of a robotics team's work from simple imitation learning to developing Vision-Language-Action (VLA) models. It outlines challenges in adapting robotic behavior to new environments, the development of models like SmolVLA and π0.5, and the shift toward simulation-generated data for training. The focus is on building robots that can see, understand, and act autonomously in dynamic settings.
Event attention history
There is not enough continuous observation data to show a trend.
Timezone · UTC
Article dates follow your selected timezone. Briefing editions use Hong Kong time (UTC+8).