Glossary · term

Multi-Scale Embodied Memory (MEM)

Multi-Scale Embodied Memory (MEM) is a memory architecture for vision-language-action robot policies. It divides memory by time scale and representation: a high-level policy maintains a compact natural-language summary of semantically important past events, while a low-level policy receives a dense window of recent observations through an efficient video encoder. MEM is a named method, not a generic label for every robot memory system.

Training2026-03-03Wave 3 · 2025–26Maturity: 3/5

Origin and context

Physical Intelligence published the MEM project page on 3 March 2026; Marcel Torne, Karl Pertsch and collaborators submitted the associated preprint the next day. Their π0.6-MEM implementation updates a language memory together with high-level subtasks and adds causal temporal attention to a vision encoder. The paper reports pretraining on robot, vision-language and video-language data, then post-training for specific robot tasks.

Sources: s1, s2

Why it matters

Long tasks create two different information problems. Recent frames can preserve motion and object location through occlusion, but retaining every image for minutes is expensive. A text summary is cheaper for facts such as which recipe step is complete, yet loses fine physical detail. MEM makes that trade-off explicit. An independent VLA review subsequently used this two-scale design as a reference point among memory-augmented policies.

Sources: s1, s3

Example

In the originating kitchen-cleanup evaluation, the long-term summary can record which objects were stored and which surfaces were cleaned, while recent visual history helps the policy continue an action after self-occlusion or change a grasp after a failed attempt. The authors evaluated policies with ten rollouts per task or recipe and tasks requiring memory for up to fifteen minutes. These are bounded study results, not evidence that MEM can safely perform arbitrary household work.

Sources: s1, s2

How it differs

Vision-Language-Action Models (VLA)

A VLA is the broader model family that maps visual and language inputs to actions. MEM is one optional architecture for adding explicit history to such a policy.

Robot Foundation Model

A robot foundation model describes a broadly reusable pretrained policy. MEM concerns memory organization and can be attached to a VLA implementation; it does not by itself make a policy foundational or broadly generalizable.

Active Context Curation

Both approaches compress history, but Active Context Curation concerns managing an agent's working context. MEM is a robot-policy design trained to combine language summaries with recent sensor observations.

Maturity and evidence

Maturity is 3. MEM has an exact, technically specified identity, independent treatment in a peer-reviewed review, and an independent research reimplementation of its short-term visual-memory mechanism. The latter used π-MEM as an experimental baseline, which shows technical uptake beyond derivative coverage. It does not replicate the language-memory component or the original long-horizon results, so maturity 4 would overstate adoption and validation.

Sources: s1, s3, s4

Limits and open questions

The complete MEM evidence still comes mainly from one author team and a March 2026 preprint. Its long-horizon evaluations use bespoke tasks, robots, data and ten rollouts per policy/task or recipe. The independent FM-VLA study reimplements only the video-memory design on a different base model and finds that force history can outperform visual memory for subtle contact events; it is not a full replication. Text summaries may omit information or propagate mistakes, and the authors identify memory beyond a single episode as future work. No reviewed source establishes production reliability or physical-safety guarantees.

Sources: s1, s3, s4

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as agent memory systems skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as computer vision skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as vision language models skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as multimodal ai skill.