Our Journey Building Vision-Language-Action Robotics
Our Journey Building Vision-Language-Action Robotics
This article details the progression of a robotics team's work from simple imitation learning to developing Vision-Language-Action (VLA) models. It outlines challenges in adapting robotic behavior to new environments, the development of models like SmolVLA and π0.5, and the shift toward simulation-generated data for training. The focus is on building robots that can see, understand, and act autonomously in dynamic settings.
This tutorial explains the evolution of a robotics team's approach from imitation learning to Vision-Language-Action (VLA) models, highlighting key milestones like SmolVLA and NVIDIA GR00T. It emphasizes the importance of data quality, model capacity, and closed-loop control in achieving adaptable robotic systems.
Source: Robotics — Paper and dataset web discovery · Read original article ↗
Article text · Machine translation into English
On this page
How we moved from repeating demonstrated motions to building robots that can see, understand and act.
One question has shaped much of our recent work in Physical AI:
How do we move from a robot repeating a demonstrated motion to actually understanding the task it is being asked to perform?
That question became especially important while experimenting with the SO-101 robotic platform. Our early imitation-learning experiments showed that a robot could successfully complete a familiar pick-and-place task when the environment closely resembled its demonstrations. But change something seemingly small, such as the starting position of the object and performance could drop significantly.
That gap between repetition and adaptation became the starting point for our journey into Vision-Language-Action models, or VLAs.
Starting With Imitation Learning
Our first baseline used ACT - Action Chunking with Transformers.
For the experiment, we trained using: 50 demonstrations 100,000 training steps Batch size: 8
When the cube started near positions well represented in the training demonstrations, the robot could complete the task. But when we changed the cube's initial position, execution became much less reliable. That observation was important. It did not mean imitation learning was ineffective.
Instead, it highlighted something fundamental about robotics:
A robot that can reproduce a trajectory is not necessarily a robot that understands the task.
And this was exactly the problem we wanted to explore next.
The Pivot: Vision + Language + Action
We began exploring Vision-Language-Action models. The idea behind VLA is powerful. Rather than training a robot only to reproduce movements, the model connects three things:
Vision - What does the robot currently see?
Language - What task has the robot been asked to perform?
Action - What movement should happen next?
Instead of learning only a fixed trajectory, the objective becomes learning a task-conditioned control policy. In our experiments, this meant combining visual observations, language instructions and continuous robot actions within a single control system.
Conceptually:
That final feedback loop is critical. Real-world robotics is not simply:
observe → predict → finish
The world changes while the robot moves. Objects shift. Grasping attempts may not land perfectly. The robot's own actions change what the cameras see. So we started treating VLA as a closed-loop Physical AI system, rather than simply evaluating an isolated AI model.
Our First Deployed VLA: SmolVLA
Our first major step was SmolVLA. At roughly 450 million parameters, the model was compact enough for us to deploy on an edge device. During our tests, we observed approximately 200 ms response latency from the VLA pipeline. Getting the model running on the robot itself was an important milestone.
For the first time, we had the complete flow operating:
Camera observation → Language instruction → VLA inference → Physical robot movement
SmolVLA performed reasonably well on short behaviors that were strongly represented in the training data. But longer tasks exposed another challenge. The robot could sometimes become trapped repeating an action instead of progressing to the next stage. We also found that performance was highly sensitive to the quality of our demonstrations. Small inconsistencies in trajectories or recordings could affect execution significantly.
This taught us another lesson:
Robotics AI is not only a model problem. It is also a data problem.
Increasing Model Capacity: π0.5
Our next experiment moved to Physical Intelligence's π0.5.
We fine-tuned the approximately 3-billion-parameter model using 100 episodes and evaluated a 5K checkpoint on the physical SO-101 robot. The change was noticeable. Compared with our SmolVLA experiments, task progression became more stable. Instructions were followed more effectively. Longer behaviours were completed with fewer repetitive actions. The larger model appeared better able to preserve the structure of a task across multiple stages. That was an encouraging result. But robotics engineering is not only about model performance. The surrounding development ecosystem matters too. Libraries, tooling, documentation, community activity and integration support can significantly influence how quickly a research experiment can evolve into a reliable engineering platform.
For our experimentation pace, we eventually decided to continue our work with another ecosystem.
Moving Toward Long-Horizon Physical AI with NVIDIA GR00T
The next stage of our VLA journey was NVIDIA GR00T. Here, the objective became more ambitious. Instead of evaluating one isolated behaviour, we started testing multiple prompts and longer task sequences. Across our own tested long-horizon scenarios, we observed an overall success rate of approximately 80%. Some of our stronger experiments involved tasks requiring several actions to remain connected as part of one larger objective, including: cleaning sequences and a ring-tower task.
In our tests, GR00T maintained task intent more consistently across multi-step execution and recovered better during longer sequences than the smaller models we had previously evaluated.
Importantly, these numbers should not be interpreted as a universal benchmark for GR00T. They reflect our own SO-101 experiments, datasets, fine-tuning configurations and task definitions. But for us, they represented a meaningful step forward.
What Changed Across the Journey?
The progression was not simply:
small model → bigger model → better robot.
Each stage exposed a different engineering problem.
ACT taught us about generalization.
The robot could imitate a known trajectory, but changes in object position revealed the limits of our dataset.
SmolVLA validated edge deployment.
We proved that vision, language and physical actions could operate together on our robotic setup. But longer tasks exposed looping behaviour and strong dependence on data quality.
π0.5 improved task structure.
Increasing model capacity produced more stable transitions between actions and better execution of longer behaviours in our experiments.
GR00T pushed us toward long-horizon Physical AI.
Instead of evaluating isolated skills, we could begin exploring how a robot maintains intent across multiple actions. And that brings us to the next major challenge.
The Next Bottleneck: Data
Training Physical AI systems requires large amounts of diverse interaction data. Collecting every possible situation manually is difficult.
Imagine trying to physically record demonstrations for every combination of:
- object position
- camera angle
- lighting condition
- robot pose
- obstacle position
- task sequence
- environment configuration
The number of combinations grows very quickly.
So we started exploring another direction: simulation-generated robotics data.
From Manual Demonstrations to Simulation
Using NVIDIA Isaac Sim and Isaac Lab, we developed reinforcement-learning environments and a pipeline capable of varying:
object placement, scene configuration and task conditions.
The resulting datasets were then evaluated on the physical SO-101 as an initial step toward simulation-to-real validation.
The idea is straightforward:
If this pipeline continues to improve, simulation could help us generate the diversity that would be extremely expensive to collect manually.
What We're Learning About Physical AI
Working through these experiments changed how we think about robotics. Building an intelligent robot is not simply a matter of choosing the largest AI model.
A working Physical AI system combines:
- Perception Can the robot understand its environment?
- Language understanding Can it connect an instruction with the current situation?
- Action generation Can it convert intent into meaningful movement?
- Temporal reasoning Can it understand that a task consists of several stages?
- Closed-loop control Can it adjust when reality does not perfectly match its expectations?
- Data quality Has the model experienced enough meaningful variations?
- Deployment engineering Can all of this run reliably on real hardware?
When these pieces begin working together, robotics starts moving beyond fixed automation.
From Programmed Robots to General-Purpose Robots
Traditional automation often works extremely well when the environment is predictable.
A robot may be programmed to:
Move here -> Pick object -> Move there -> Place object
But change the object's position, introduce another object or modify the task, and the workflow may require reprogramming. Physical AI points toward something different.
Instead of programming every movement, we want robots that can receive an instruction such as:
"Pick up the red object and place it inside the container."
The system then needs to determine:
- What is the red object?
- Where is it?
- Where is the container?
- How should I grasp it?
- How should I move?
- Did my grasp succeed?
- What should I do next?
That shift from executing predefined motion to interpreting task intent is what makes Vision-Language-Action research particularly exciting for us.
Where We Go Next
Our benchmark is still evolving. We are continuing to evaluate additional open-source VLA models across simple manipulation, long-horizon tasks, generalization, multi-task behaviour, real-world deployment and simulation-to-real transfer.
The objective is not simply to find a model that performs well on one demonstration. The larger goal is to develop a reliable multi-task robotic policy and a repeatable way of measuring how well it generalizes. Our journey started with a robot learning to repeat a motion.
The direction now is much broader - Building robots that can see, understand, adapt and act.
And we're still learning with every experiment.
What the source reports
Publisher-reported claims, with original evidence. These results have not been independently verified by RoboSignal.
Reported numbers
training steps
100,000
demonstrations
50 episodes
parameters
450,000,000
response latency
≈200
episodes
100 episodes
checkpoints
5,000
Source excerpts and review record
Automatically extracted; no manual editorial approval recorded.
ation Learning Our first baseline used ACT - Action Chunking with Transformers . For the experiment, we trained using: 50 demonstrations 100,000 training steps Batch size: 8 When the cube started near positions well represented in the training demonstrations, the robot could complete the task. But when we changed the cube's initial position, execution became much less reliable. That observation was important. It did
Open source S4
r First Deployed VLA: SmolVLA Our first major step was SmolVLA . At roughly 450 million parameters , the model was compact enough for us to deploy on an edge device. During our tests, we observed approximately 200 ms response latency from the VLA pipeline. Getting the model running on the robot itself was an important milestone. For the first time, we had the complete flow operating: Camera observation → Language ins
Open source S8
ve became more ambitious. Instead of evaluating one isolated behaviour, we started testing multiple prompts and longer task sequences. Across our own tested long-horizon scenarios, we observed an overall success rate of approximately 80% . Some of our stronger experiments involved tasks requiring several actions to remain connected as part of one larger objective, including: cleaning sequences and a ring-tower task.
Open source S13
rajectories or recordings could affect execution significantly. This taught us another lesson: Robotics AI is not only a model problem. It is also a data problem. Increasing Model Capacity: π0.5 Our next experiment moved to Physical Intelligence's π0.5 . We fine-tuned the approximately 3-billion-parameter model using 100 episodes and evaluated a 5K checkpoint on the physical SO-101 robot. The change was noticeable. C
Open source S10
Source:Robotics — Paper and dataset web discovery · linkedin.com