Reinforcement Fine-Tuning (RFT)
Reinforcement fine-tuning, or RFT, is post-training in which a model samples responses, a grader or reward function scores them, and optimization increases the probability of higher-reward behavior. It adapts a pretrained model to a target task without requiring one prescribed answer for every prompt. RFT is broader than reinforcement learning from verifiable rewards, which restricts the signal to outcomes that can be checked automatically, and broader than any single optimizer such as GRPO.
Origin and context
A November 2023 paper used reinforcement-learning fine-tuning for a language-model training phase based on human or AI feedback and studied its inductive biases. OpenAI then presented Reinforcement Fine-Tuning in December 2024 as a productized technique for verifiable, domain-specific work. By December 2025, AWS offered RFT in Amazon Bedrock with rule-based or AI-based graders. A separate 2025 position paper used RFT as a broader umbrella for reward-driven reasoning improvements. The category therefore spans research vocabulary and product workflows rather than one fixed API.
Why it matters
Many domain tasks have outputs that are easy to score but expensive to demonstrate perfectly. A code test, mathematical verifier, structured rule, or model judge can evaluate several attempted solutions and provide a learning signal. That makes RFT attractive when teams can define success more reliably than they can write ideal completions. The difficult part shifts to grader design: a reward function can be incomplete, exploitable, or misaligned with the quality users actually need.
Example
A team adapting a model for structured data extraction supplies prompts and a grader that checks schema validity and selected field-level rules. During training, the model generates multiple candidate outputs; valid and more accurate candidates receive higher scores, and the policy is updated accordingly. If the grader checks only JSON syntax, the model may learn to emit well-formed but incorrect records. Human-held validation data and adversarial tests are therefore part of evaluating the trained model, even when the training signal is automated.
How it differs
RLHF
RLHF is a family of workflows grounded in human preference feedback, often through a learned reward model. RFT describes reward-driven task customization more broadly and can use deterministic graders, model judges, or other signals without collecting pairwise human preferences.
Self-Rewarding Models (SRM)
A self-rewarding model generates or judges its own supervision. RFT does not specify who supplies the reward: the grader may be external, rule-based, human-derived, or another model.
Maturity and evidence
Maturity is rated 4. The label has documented offerings from two independent cloud-model providers and broad research usage across language and multimodal reasoning. The implementation remains less standardized than the name: providers expose different graders, supported models, optimization details, and evaluation practices.
Limits and open questions
RFT can optimize a proxy rather than the intended task or exploit weaknesses in a grader. The reviewed 2023 experiment found that reinforcement-learning fine-tuning favored more extractable features in its controlled settings, with implications for robustness and generalization. The 2025 position paper identifies reward hacking as a central challenge for RFT. These findings remain setting-specific, so results should be reported with the model, grader, training distribution, and evaluation used.
Related terms
References
- 12 Days of OpenAI: Reinforcement Fine-TuningOpenAI · 2024-12-06 · class A
- Amazon Bedrock now supports reinforcement fine-tuningAmazon Web Services · 2025-12-03 · class A
- Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language ModelsIndependent research team / arXiv · 2025-05-24 · class A
- Reinforcement Learning Fine-tuning of Language Models is Biased Towards More Extractable FeaturesAI Safety Hub Labs / arXiv · 2023-11-07 · class A
- Training language models to follow instructions with human feedbackOpenAI / arXiv · 2022-03-04 · class A
- Self-Rewarding Language ModelsMeta and New York University / arXiv · 2024-01-18 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as reinforcement learning skill.
This term is also covered in the Skills Atlas as model training skill.
This term is also covered in the Skills Atlas as reward modeling skill.