Vision-Language-Action Models (VLA)
A vision-language-action (VLA) model is a multimodal policy that conditions on visual observations and language instructions and produces actions for an embodied system, usually a robot. It connects perception and language grounding to control in one learned model or tightly integrated policy. A vision-language model that only describes an image is not a VLA; the defining output must represent an action or control decision.
Origin and context
The 2023 RT-2 paper used vision-language-action models as a category name and represented robot actions as tokens alongside language. OpenVLA followed in 2024 with an open 7-billion-parameter model trained on a reported 970,000 robot demonstrations, plus checkpoints and adaptation tools. The π0 work, first submitted as a preprint in October 2024 and later published at RSS 2025, instead used a flow-matching action architecture on top of a pretrained vision-language model. These examples establish a family, not a requirement to encode every action as text.
Why it matters
VLA models seek to reuse broad visual and language representations while learning physical action, reducing the need to build an isolated policy for every instruction and object. RT-2 reported improved generalization to novel objects and commands in its evaluations. OpenVLA made checkpoints and training tools available for adaptation, and π0 addressed continuous, dexterous control with a different action-generation method. This line of work matters for general-purpose robotics, but results from selected tasks and laboratory setups do not establish dependable operation in uncontrolled environments.
Example
In a Skills Intelligence illustration, a robot receives camera images and the instruction “put the red cup in the sink.” A VLA policy uses the scene and language to produce control outputs, such as end-effector movements. RT-2 represents actions through discrete tokens; π0 uses flow matching for continuous action generation. The implementation difference matters when adapting to a robot's action space. A system that only captions the scene is a vision-language model, while a generalist robot policy is not necessarily language-conditioned and therefore is not automatically a VLA.
Maturity and evidence
Skills Intelligence rates VLA models at maturity 3. RT-2, OpenVLA, and π0 demonstrate distinct architectures and adaptation approaches, while a separate survey organizes a broader research literature under the same category. Some foundational papers share collaborators, so they should not be counted as wholly independent validation of one another. Research use is established; comparable evidence for dependable deployment across uncontrolled environments remains a different and more demanding test.
Limits and open questions
Robot demonstrations are expensive and uneven, action spaces differ across hardware, and small perception errors can become physical failures. Reported success rates depend on task definitions, embodiments, training data, and laboratory conditions, which complicates comparison. Internet-derived semantics can also be poorly grounded in a specific robot's capabilities. VLA is a broad architecture category, not evidence that a system can safely generalize to arbitrary instructions, objects, people, or environments.
Related terms
References
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlGoogle DeepMind / arXiv · 2023-07-28 · class A
- OpenVLA: An Open-Source Vision-Language-Action Model (arXiv v3)OpenVLA collaboration / arXiv · 2024-09-05 · class A
- π0: A Vision-Language-Action Flow Model for General Robot Control (RSS 2025; arXiv v4)Physical Intelligence / arXiv · 2026-01-08 · class A
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges (preprint, v2)Sapkota et al. / arXiv · 2026-01-29 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as multimodal ai skill.
This term is also covered in the Skills Atlas as computer vision skill.
This term is also covered in the Skills Atlas as reinforcement learning skill.
This term is also covered in the Skills Atlas as vision language models skill.