Glossary · term

Robot Foundation Model

A robot foundation model is a broadly pretrained model for robot behavior that can be reused across multiple tasks, environments or embodiments and adapted to new settings. Its defining intent is transfer and adaptation rather than one fixed policy for one robot-task pair. Inputs and outputs can vary by system: some models map vision and language to actions, while others generate navigation subgoals, policies or intermediate representations. The term therefore does not imply that every model directly emits low-level motor commands.

Training2023-06-26Wave 3 · 2025–26Maturity: 3/5

Origin and context

Version 1 of ViNT used the robot-foundation-model category in June 2023; version 2 added an explicit definition based on zero-shot deployment in useful novel settings and adaptation to downstream tasks. The work was accepted at CoRL 2023. A short paper accepted to the RLC 2024 Workshop on Training Agents with Foundation Models applied the category to task-specific policy generation. NVIDIA's 2025 GR00T N1 technical-report preprint describes a generalist humanoid foundation model implemented as a vision-language-action system. These sources span different teams and robot settings, but later evidence remains workshop- or preprint-stage.

Sources: s1, s2, s3

Why it matters

Robot learning is often limited by data tied to one embodiment, environment or task. A reusable pretrained model can provide representations or policies that reduce the amount of task-specific training and make cross-platform adaptation a measurable research goal. The category also gives teams a way to ask whether a model is genuinely transferable or merely large. It does not erase embodiment differences: sensors, action spaces, timing, dynamics and safety constraints still require explicit interfaces, adaptation and physical validation.

Sources: s1, s2, s3

Example

A navigation model trained across several robots and environments may accept camera observations and a goal, deploy zero-shot on a new route, and later be fine-tuned for a different platform. That can qualify as a robot foundation model even if it predicts waypoints rather than joint torques. A vision-language-action model trained for one arm and one narrow benchmark is not automatically a robot foundation model: the VLA interface describes modalities and actions, whereas the foundation-model claim depends on demonstrated breadth and adaptation.

Sources: s1, s3

How it differs

Vision-Language-Action Models (VLA)

VLA names an architectural input-output pattern connecting vision and language to actions. A robot foundation model names a transfer and reuse role. A system such as GR00T N1 can be both, but neither category logically contains every instance of the other.

World Foundation Model

A world foundation model predicts or generates environment states and can support simulation or planning. A robot foundation model centers reusable robot behavior or policy. A robotic system may combine both layers without making the terms synonyms.

Maturity and evidence

Maturity is rated 3. The category has an explicit definition in the October 2023 ViNT revision, a peer-reviewed originating paper, independent follow-up and multiple model families across navigation, policy generation and humanoid control. It remains below 4 because evaluation protocols for cross-task and cross-embodiment generality are not standardized, and broad claims often depend on simulations or demonstrations from the proposing organization.

Sources: s1, s2, s3

Limits and open questions

Foundation-model branding does not itself demonstrate robust transfer. Results can depend on proprietary data, embodiment-specific adapters and benchmark choices, while simulation performance may not transfer safely to physical hardware. Reports should separate zero-shot deployment, fine-tuning and hardware adaptation and state the sensors, action representation and tested embodiments. A model's breadth should be supported by evaluations rather than inferred from parameter count or the word 'generalist'.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as multimodal ai skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as vision language models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.