RL Token (RLT)
RL Token (RLT) is a two-stage method for adapting a vision-language-action model with online reinforcement learning. First, an encoder-decoder is trained to compress the VLA's internal embeddings into a learned bottleneck representation called the RL token. The feature model is then frozen, and lightweight actor and critic networks use that representation, robot state and the VLA's reference action chunk to learn task-specific refinements. The token is therefore an interface to a policy head, not a complete policy by itself.
Origin and context
Physical Intelligence published the RLT project on 19 March 2026 and posted the paper to arXiv on 24 April. The authors positioned it as a way to improve the precise, contact-rich phase of a broader manipulation behavior without updating a full VLA during online practice. Their experiments used four real-robot tasks: screw installation, zip-tie fastening, Ethernet insertion and power-cord insertion. Later independent projects implemented the method on OpenPI, Franka and ManiSkill paths.
Why it matters
RLT separates broad pretrained perception and action proposals from a small component that can be updated from limited robot interaction. That architecture makes the sample-efficiency question concrete: instead of relearning an entire task, the system can target an insertion or alignment phase. The authors report faster and more successful execution in their four setups, including one run with 15 minutes of collected robot data over two hours of wall-clock training. Independent codebases make the method inspectable, but do not yet reproduce those headline results.
Example
In a controlled research cell, a team could train the RLT bottleneck on demonstrations from a supported VLA, freeze that feature stack, and let a small actor-critic refine only the final cable-insertion phase. Evaluation should compare a fixed base policy and RLT under the same robot, cameras, reward, reset procedure and action timing. Because online exploration moves hardware, researchers need physical barriers, force and workspace limits, emergency stops, human supervision and a separate validation set before any deployment claim.
How it differs
Vision-Language-Action Models (VLA)
A VLA is the broader model architecture that maps perception and language to actions. RLT is an adaptation method layered on a VLA and does not replace that backbone.
Robot Foundation Model
Robot foundation model describes a pretrained policy family. An RL token is a compact interface used to specialize one such model for a narrow phase through online learning.
Reinforcement Fine-Tuning (RFT)
Reinforcement fine-tuning is a broad family of reward-driven adaptation methods. RLT specifies a particular bottleneck representation, frozen VLA feature stack and lightweight actor-critic design for robot actions.
Physical AI
Physical AI is an umbrella category for systems that perceive and act in the world. RLT is one concrete robot-learning technique within that broader domain.
Maturity and evidence
Maturity is 3. RLT has a named primary paper and project page, a complete technical recipe, real-robot experiments, and multiple independent open implementations with executable configurations. It is no longer merely a proposed label. It is not rated higher because the paper is recent and not yet peer reviewed, the original training code and task data are not a complete public reproduction package, and no independent team located in this review has replicated the four reported hardware results.
Limits and open questions
The reported gains are study-bounded: four manipulation tasks, one main VLA family, specific cameras, action chunks, rewards, interventions and hardware. Minutes of robot data are not the same as total wall-clock time or engineering effort, and faster execution is not a general safety or robustness result. Independent implementations have not reproduced the headline comparisons, and one RLinf issue questions whether its decoder matches the paper's reconstruction objective. The learned token may omit information needed outside its training distribution, while online exploration can damage equipment or create unsafe motion. Any production use requires new task-level validation and physical safety review.
Related terms
References
- Precise Manipulation with Efficient Online RLPhysical Intelligence · 2026-03-19 · class A
- RL Token: Bootstrapping Online RL with Vision-Language-Action ModelsPhysical Intelligence · 2026-03-19 · class A
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models — arXiv recordarXiv · 2026-04-24 · class A
- RL Token: Bootstrapping Online RL with Vision-Language-Action Models — RLinf documentationRLinf · 2026 · class B
- openpi-RLT: Real-Robot RLT Reproduction on OpenPIYi Yang, Huaihang Zheng, Kai Ma and collaborators · 2026 · class B
- Potential discrepancy between the RLT decoder and Equation 2 of the paperRLinf GitHub community · 2026-07-17 · class C
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as reinforcement learning skill.
This term is also covered in the Skills Atlas as vision language models skill.
This term is also covered in the Skills Atlas as model training skill.
This term is also covered in the Skills Atlas as computer vision skill.
This term is also covered in the Skills Atlas as fine tuning evaluation skill.