Glossary · term
Test-time RL
A method (TTRL) that applies reinforcement learning to unlabeled data at inference time, allowing a model to improve itself without ground-truth labels. It uses majority voting across multiple responses as a reward signal. Work by Yuxin Zuo et al. (Tsinghua/Shanghai AI Lab, April 2025).
Training2025–2026Wave 2 · 2024Maturity: 1/5
Maturity rationale
single R2 source, 2025-26 neologism
References
- Zuo et al. 2025 — TTRL(arxiv)
Author: Społeczność / Anonimowi