Spatial Intelligence
Spatial intelligence is the capacity to acquire, represent, transform and reason about spatial information such as position, distance, direction, shape, viewpoint and motion, then use it to solve problems, predict change or guide action. In AI, it spans perception, memory, reasoning, generation, navigation and manipulation. It is a capability family, not a single architecture, benchmark, world model or World Labs product.
Origin and context
The exact phrase is not a 2025 coinage. Gardner used Spatial Intelligence in Frames of Mind in 1983, while empirical spatial-ability research is older. A 2006 National Research Council synthesis treated spatial intelligence as one of several overlapping labels and analyzed spatial thinking through concepts, representations and reasoning. Fei-Fei Li's May 2024 TED talk later applied the phrase to AI that processes visual data, predicts and acts. World Labs subsequently adopted it as a company-wide framing around models that perceive, generate, reason and interact.
Why it matters
Spatial tasks require maintaining geometry across viewpoints and time, not merely naming visible objects. That matters for scene understanding, navigation, manipulation, autonomous systems, generated environments and planning. A 2026 survey of vision-language models organizes evidence across spatial perception, understanding and extrapolation, while reporting divergent results and benchmark-design biases. Independent embodied-agent research illustrates another slice: structured landmarks, routes and map-like memory for navigation. Separating these competencies lets teams test a specific failure mode instead of claiming one undifferentiated, human-like intelligence.
Example
A model may correctly say that a cup is left of a plate in one image yet fail after the camera moves, confuse viewer-relative and map-relative directions, or lose a route after several turns. A fuller evaluation would separate relation recognition, viewpoint transformation, metric estimation, memory, prediction, planning and grounded action. Conversely, World Labs' Marble can generate an explorable 3D scene, but coherent-looking output alone does not demonstrate reliable physics, long-horizon memory, navigation or robot control.
How it differs
World Models
A world model learns a representation or predictor of an environment for imagining futures or control. It can support spatial intelligence, but neither it nor a foundation-scale variant proves the broader capability; visual generation alone is insufficient evidence of spatial reasoning.
Physical AI
Physical AI names a whole AI-enabled physical system and its sensing-and-action loop. Spatial intelligence is one capability such a system may require and can also be studied in images or virtual environments, so the terms are not synonyms.
Vision-Language-Action Models (VLA)
VLA names a model or policy that conditions on vision and language and emits actions. Those modalities do not guarantee robust spatial representation, memory or geometry, and systems without a VLA can still solve spatial tasks.
Maturity and evidence
Skills Intelligence rates the term at maturity 3 with an established lifecycle. The label and its research tradition predate modern generative AI, and independent peer-reviewed studies now apply it to both vision-language models and embodied agents. It remains below 4 because human taxonomies differ, AI terminology is inconsistent and no standard test covers the whole capability. Convergent metrics and reproducible transfer from static tests to navigation and manipulation would justify a higher rating.
Limits and open questions
Do not infer general spatial intelligence from one benchmark, attractive 3D output or a successful robot demo. Embodied AI describes an agent situated in and acting through an environment; embodiment does not guarantee broad spatial reasoning. Spatial computing describes technologies and interaction organized around location, physical or virtual space and spatial data, not a cognitive score. World Labs' roadmap is one commercial interpretation whose products still require explicit tests of geometry, dynamics, persistence and action.
Related terms
References
- Frames of Mind: The Theory of Multiple IntelligencesBasic Books / Google Books · 1983-11-23 · class A
- Learning to Think SpatiallyNational Research Council / National Academies Press · 2006 · class A
- A Heuristic Framework of Spatial Ability: a Review and Synthesis of Spatial Factor Literature to Support its Translation into STEM EducationEducational Psychology Review · 2018-03-02 · class A
- With spatial intelligence, AI will understand the real worldTED · 2024-05-16 · class A
- About World LabsWorld Labs · 2026 · class A
- Spatial intelligence in vision-language models: a comprehensive surveyArtificial Intelligence Review · 2026-08-11 · class A
- Brain-inspired spatial intelligence for embodied agentsNature Communications · 2026-06-27 · class A
- World ModelsDavid Ha and Jürgen Schmidhuber / arXiv · 2018-03-27 · class A
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic ControlGoogle DeepMind and Everyday Robots / arXiv · 2023-07-28 · class A
- Spatial ComputingThe MIT Press · 2020-02-18 · class A
- Inside Fei-Fei Li’s Plan to Build AI-Powered Virtual WorldsTIME · 2025-12-09 · class B
- Physical AI and Data Generation for RoboticsNational Institute of Standards and Technology · 2018-12-11 · class A
Last updated: 2026-09-07