Glossary · term

World Foundation Model

A world foundation model, or WFM, is a broadly pretrained world model intended to predict or generate physical-world states and to be adapted for multiple downstream simulation, planning or physical-AI tasks. Many current examples operate on video and can condition generation on text, images, actions or other state signals. The foundation-model qualifier adds reusable pretraining and adaptation to the broader world-model idea; it does not guarantee physically correct simulation or require one particular modality or vendor platform.

Training2025-01-06Wave 3 · 2025–26Maturity: 3/5

Origin and context

NVIDIA announced the Cosmos World Foundation Model platform on 6 January 2025 and submitted its technical-report preprint the following day. A COLM 2025 study later investigated test-time scaling for world simulation, although it includes an NVIDIA coauthor. Independently, Xiaomi Robotics' July 2026 U0 preprint applies the category to reusable generation adapted for embodied tasks, using EMU3.5 initialization. That is independent research usage of the term, not an independent verification of either vendor's performance claims.

Sources: s1, s2, s3, s4

Why it matters

Physical systems need data about how environments change, including rare or expensive situations that are difficult to collect with real hardware. A reusable pretrained world model can generate candidate trajectories, support simulation, provide synthetic training data or help a planner compare possible futures. The category separates that environment-prediction layer from a robot policy that chooses actions. Its practical value depends on fidelity: visually plausible video may still violate geometry, contact dynamics or causal effects and therefore cannot be assumed safe or accurate for deployment.

Sources: s1, s2, s3, s4

Example

A WFM can take recent camera frames plus a proposed control signal and generate possible future frames for a driving or robotics simulator. A downstream team might adapt the pretrained model to a particular factory layout and use those rollouts during policy development. A text-to-video model that produces attractive scenes without representing action-conditioned environment evolution is not automatically a WFM. Conversely, a compact task-specific dynamics model may be a world model but not a foundation model if it lacks broad reusable pretraining.

Sources: s2, s3

How it differs

World Models

World models are the broader class of learned representations or predictors of environment dynamics. A WFM is a pretrained, reusable member of that class designed for adaptation across multiple downstream domains or tasks.

Robot Foundation Model

A robot foundation model centers transferable robot behavior or policy. A WFM centers prediction or generation of environment states. They can be combined in a planner, but predicting a future does not by itself select or execute an action.

Maturity and evidence

Maturity is rated 3. The category now has a dated public framing, a detailed Cosmos technical preprint, peer-reviewed research and independent usage in Xiaomi Robotics' U0 preprint. It is not confined to one model family. The rating stays below 4 because reusable-world-model research remains recent, terminology overlaps with video foundation models, and claimed benefits depend on the downstream task. Xiaomi's preprint is evidence of independent usage, not peer-reviewed validation.

Sources: s1, s2, s3, s4

Limits and open questions

Realistic frames are not proof of accurate environment dynamics. Xiaomi's preprint specifically distinguishes visual generation from the geometric, multi-view and embodiment constraints of robotics. Evaluation therefore needs to examine what the generated states support in the intended task, rather than relying only on attractive demonstrations. Vendor-reported improvements and research experiments have bounded conditions; this glossary does not infer operational reliability from either.

Sources: s2, s3, s4

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as multimodal ai skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as video generation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.