Skip to content

What is a vision-language-action (VLA) model?

RoboSignal · Published · Version 1.0

A vision-language-action model connects visual observations and language instructions to robot actions. The important distinction is its action output: a description of what a robot should do is not the same as a command the robot can execute.

From instruction to action

OpenVLA is a concrete VLA example. Its project describes a model for robot manipulation and publishes evaluation and adaptation information. Use the project and paper to inspect the supported robot setup, observation inputs and action representation. A VLA label alone does not specify control frequency, safety mechanisms or deployment readiness.

What to check in a VLA announcement

Record the model version, robot platform, camera configuration and action space. Ask whether a result comes from the released checkpoint or from a checkpoint fine-tuned on the target task. Note how success was measured, the number of attempts and the role of human resets. Keep simulation results separate from physical trials. A task score on one setup does not answer whether a different robot can use the same policy.

VLA models and foundation models

These terms describe different properties. VLA describes the connection between modalities and actions. Foundation model describes a model intended for adaptation across downstream uses. A model may fit both descriptions, but neither term establishes general physical competence. Compare a release through its documented tasks and transfer tests, rather than through the breadth of its name. See the robot foundation model guide for the broader context.

Primary references

Reference links checked 2026-10-04. Project claims remain attributed to their original source. This page is not a certification or a live test of the referenced system.

Related reading

Cite this reference

RoboSignal. “What is a vision-language-action (VLA) model?” (2026-10-04), version 1.0. https://robosignal.ai/glossary/vision-language-action-model. Cite the original project separately for its own reported results.

Corrections and feedback · Reference RSS