Multi-Scale Embodied Memory (MEM)
Multi-Scale Embodied Memory (MEM) is a memory architecture for vision-language-action robot policies. It divides memory by time scale and representation: a high-level policy maintains a compact natural-language summary of semantically important past events, while a low-level policy receives a dense window of recent observations through an efficient video encoder. MEM is a named method, not a generic label for every robot memory system.
Origin and context
Physical Intelligence published the MEM project page on 3 March 2026; Marcel Torne, Karl Pertsch and collaborators submitted the associated preprint the next day. Their π0.6-MEM implementation updates a language memory together with high-level subtasks and adds causal temporal attention to a vision encoder. The paper reports pretraining on robot, vision-language and video-language data, then post-training for specific robot tasks.
Why it matters
Long tasks create two different information problems. Recent frames can preserve motion and object location through occlusion, but retaining every image for minutes is expensive. A text summary is cheaper for facts such as which recipe step is complete, yet loses fine physical detail. MEM makes that trade-off explicit. An independent VLA review subsequently used this two-scale design as a reference point among memory-augmented policies.
Example
In the originating kitchen-cleanup evaluation, the long-term summary can record which objects were stored and which surfaces were cleaned, while recent visual history helps the policy continue an action after self-occlusion or change a grasp after a failed attempt. The authors evaluated policies with ten rollouts per task or recipe and tasks requiring memory for up to fifteen minutes. These are bounded study results, not evidence that MEM can safely perform arbitrary household work.
How it differs
Vision-Language-Action Models (VLA)
A VLA is the broader model family that maps visual and language inputs to actions. MEM is one optional architecture for adding explicit history to such a policy.
Robot Foundation Model
A robot foundation model describes a broadly reusable pretrained policy. MEM concerns memory organization and can be attached to a VLA implementation; it does not by itself make a policy foundational or broadly generalizable.
Active Context Curation
Both approaches compress history, but Active Context Curation concerns managing an agent's working context. MEM is a robot-policy design trained to combine language summaries with recent sensor observations.
Maturity and evidence
Maturity is 3. MEM has an exact, technically specified identity, independent treatment in a peer-reviewed review, and an independent research reimplementation of its short-term visual-memory mechanism. The latter used π-MEM as an experimental baseline, which shows technical uptake beyond derivative coverage. It does not replicate the language-memory component or the original long-horizon results, so maturity 4 would overstate adoption and validation.
Limits and open questions
The complete MEM evidence still comes mainly from one author team and a March 2026 preprint. Its long-horizon evaluations use bespoke tasks, robots, data and ten rollouts per policy/task or recipe. The independent FM-VLA study reimplements only the video-memory design on a different base model and finds that force history can outperform visual memory for subtle contact events; it is not a full replication. Text summaries may omit information or propagate mistakes, and the authors identify memory beyond a single episode as future work. No reviewed source establishes production reliability or physical-safety guarantees.
Related terms
References
- MEM: Multi-Scale Embodied Memory for Vision Language Action Models, version 2Torne et al. / arXiv · 2026-03-08 · class A
- VLAs with Long and Short-Term MemoryPhysical Intelligence · 2026-03-03 · class A
- Vision-Language-Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A ReviewDrones / MDPI · 2026-05-26 · class B
- FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich ManipulationLi et al. / arXiv · 2026-07-20 · class B
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as agent memory systems skill.
This term is also covered in the Skills Atlas as computer vision skill.
This term is also covered in the Skills Atlas as vision language models skill.
This term is also covered in the Skills Atlas as model training skill.
This term is also covered in the Skills Atlas as multimodal ai skill.