Glossary · term

Physical AI

Physical AI is an umbrella category for AI-enabled systems that sense a physical environment, make decisions and produce actions through a machine, robot or other embodied platform. It describes a system domain rather than a single model architecture, training method or vendor stack. A vision-language-action model, robot foundation model or world foundation model can contribute to such a system, but none is synonymous with the category.

Other2018-12-11Wave 3 · 2025–26Maturity: 3/5

Origin and context

The exact label predates NVIDIA's recent campaign. NIST created its 'Physical AI and Data Generation for Robotics' project page in December 2018, using the term for measurement and deployment of AI-enhanced manufacturing robots. In 2020, Miriyev and Kovač defined 'physical artificial intelligence' more narrowly as the theory and practice of synthesizing nature-like intelligent robotic systems, emphasizing the joint design of body, control, sensing and actuation. NVIDIA's January 2025 Cosmos announcement later promoted a broader commercial category spanning robots and autonomous vehicles. The European Commission's 2026 challenge independently adopted that broader umbrella in an official funding programme.

Sources: s1, s2, s3, s4

Why it matters

The label is useful when the unit of analysis is the whole closed-loop physical system: sensors, learned representations, decision or control models, actuators, data pipelines, simulation, deployment constraints and evaluation. That scope helps teams avoid treating a strong model benchmark as evidence that a machine will work safely or reliably outside the lab. NIST's programme focuses on metrics and test methods for AI-enabled robots, while the EU challenge requires prototypes, access to real-world testing and interest from end users or integrators. Those requirements illustrate why physical deployment, not model branding, is the practical boundary.

Sources: s1, s4

Example

Consider a warehouse mobile manipulator that uses cameras to locate a package, plans a route, adjusts its grasp from sensor feedback and executes the task under an operating policy. The complete loop is a physical-AI system. Its VLA controller describes one perception-language-action interface; a robot foundation model describes a reusable behavior model; a WFM may generate candidate future states for simulation. A chatbot that only advises an operator is not physical AI in this sense because it does not close the sensing-and-action loop through a physical system.

Sources: s1, s3, s4

How it differs

World Foundation Model

A world foundation model predicts or generates environment states for reuse across tasks. It may supply simulation or training data to a physical-AI system, but it does not by itself provide the sensors, control loop, actuators or deployed machine that make the broader system physical.

Vision-Language-Action Models (VLA)

A VLA model names a perception-language-action interface or policy architecture. Physical AI names the wider application and system domain. A VLA can be one component of a physical-AI system, while physical systems can use other control architectures.

Maturity and evidence

Skills Intelligence rates the term at maturity 3 with an established lifecycle. The exact label has dated use since at least 2018, a formal academic framing from 2020, and independent institutional adoption in the 2025-26 NVIDIA and EU materials. It remains below 4 because meanings range from morphology-and-control co-design to a broad commercial robotics stack, and no reviewed source establishes a standard boundary or general operational reliability. A stable cross-organization taxonomy and comparable deployment evidence would support a higher rating.

Sources: s1, s2, s3, s4

Limits and open questions

Physical AI is not evidence that a system understands physics, adapts robustly or operates safely. Its boundary with embodied AI is not standardized: the 2020 formulation emphasizes intelligence emerging from the co-design of body and control, whereas current umbrella usage can include conventional hardware combined with learned models and simulation. NVIDIA's association of the category with Cosmos and its robotics stack documents a vendor framing, not a requirement for those products or proof of physical fidelity. Evaluations should name the sensors, actions, hardware, environment and operating conditions instead of relying on the label.

Sources: s1, s2, s3, s4

Related terms

References

Last updated: 2026-09-05