World Models
A world model is a learned representation that predicts relevant aspects of an environment and how they may change under actions. An agent can use those predictions to evaluate possible futures, learn a policy from imagined experience, or construct useful internal state. The term describes a functional role rather than one architecture: a world model may predict observations, latent states, rewards, or other task-relevant quantities.
Origin and context
Ha and Schmidhuber's 2018 World Models paper trained a compressed visual representation and a recurrent dynamics model, then optimized a small controller using the learned environment. LeCun's 2022 position paper placed a configurable predictive world model inside a proposed architecture for autonomous intelligence, while explicitly presenting that design as a research path. DreamerV3 provided an independent 2023 demonstration of learning behavior by imagining future scenarios in a world model across more than 150 reported tasks. These works use related ideas without defining one canonical implementation.
Why it matters
World models can let an agent learn or plan from internal predictions instead of relying only on direct trial and error in the real environment. That is attractive when real interactions are slow, costly, or risky, and when useful representations must capture change over time. DreamerV3's cross-domain experiments show why the approach matters for reinforcement learning, while LeCun's proposal illustrates its broader role in research on planning and hierarchical prediction. Neither result establishes that a learned model contains a complete, human-like understanding of physical reality.
Example
In a simulated driving task, a world model could encode the current scene into a latent state and predict how that state, along with a reward signal, changes after steering or braking. A controller can compare imagined action sequences before choosing one. In the 2018 study, a controller was trained inside generated rollouts and transferred back to the environment. A video generator that produces plausible clips is not automatically an agent world model: the label requires evidence that its predictions support state estimation, planning, control, or another specified model-based function.
Maturity and evidence
World models merit maturity 3. The concept has multiple detailed formulations and demonstrated reinforcement-learning systems from independent teams, so it is more than a speculative label. Implementations and evaluation criteria remain heterogeneous, and influential proposals still frame key capabilities as future research. Evidence of reliable transfer, calibrated long-horizon prediction, and comparable evaluation across real-world domains would support a higher rating.
Limits and open questions
A learned model can omit rare events, compound small prediction errors, or represent only what helps its training objective. Planning can then exploit inaccuracies rather than produce valid behavior in the real environment. The phrase world model is also used loosely across reinforcement learning, robotics, video generation, and cognitive speculation. Editors should identify the predicted variables and intended use instead of inferring physical understanding from visual coherence or from the label alone.
Related terms
References
- World ModelsDavid Ha and Jürgen Schmidhuber / arXiv · 2018-03-27 · class A
- A Path Towards Autonomous Machine IntelligenceOpenReview · 2022-06-27 · class A
- Mastering Diverse Domains through World ModelsGoogle DeepMind / arXiv · 2023-01-10 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as reinforcement learning skill.
This term is also covered in the Skills Atlas as deep learning skill.