Glossary · term

Post-training

Post-training is the set of weight-updating stages applied after a foundation model's broad pretraining. For language models it commonly includes supervised instruction tuning, preference optimization such as DPO or RLHF, reinforcement learning with verifiable rewards, and targeted safety or capability training. It is a phase of model development, not one fixed algorithm, and it is distinct from prompting or retrieval performed only at inference time.

Training2019-04-03Wave 1 · 2023Maturity: 3/5

Origin and context

Xu and colleagues used BERT post-training in 2019 for domain adaptation before task fine-tuning. Ke and colleagues used posttraining in 2022 for continual adaptation to unlabeled domain corpora. OpenAI's 2023 GPT-4 materials then contrasted pretraining with a behavior-shaping post-training process, while Tulu 3 in 2024 published a reproducible multi-stage recipe spanning supervised fine-tuning, DPO, and RLVR. These uses document an expanding scope without establishing a single originator.

Sources: s5, s4, s1, s2

Why it matters

Pretraining produces a model that predicts likely continuations; post-training can make that base model follow instructions, prefer useful responses, acquire specialized behaviors, or comply more reliably with a product's policies. The phase therefore strongly affects the behavior users experience. It also concentrates difficult choices about training data, reward signals, evaluators, regressions, and trade-offs between helpfulness, safety, calibration, and retained capabilities.

Sources: s1, s2, s3

Example

A team may start with a pretrained language model, run supervised fine-tuning on demonstrations, optimize it on chosen-versus-rejected answers, and finally use rule-checkable tasks for reinforcement learning. Those stages together form a post-training recipe. Serving the resulting model with a longer prompt or a retrieval system changes its inputs at runtime and is not, by itself, post-training.

Sources: s2, s3

Maturity and evidence

Maturity is rated 3. Multiple independent organizations use the term, and Tulu 3 provides an open implementation and evaluation record for a multi-stage recipe. The label is established but not standardized: organizations draw its boundary differently, individual methods evolve quickly, and public evidence rarely reveals the complete proprietary pipeline. Those variations make a stronger maturity claim premature.

Sources: s1, s2, s3

Limits and open questions

Post-training does not guarantee alignment, factuality, or durable capability gains. Outcomes depend on the base model, data coverage, reward design, sampling policy, and evaluation protocol. A method can improve one benchmark while harming calibration or another behavior, and a published recipe may not transfer to a different model family. Reports should name the exact stages and datasets instead of using post-training as an unexplained catch-all.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm fine tuning skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as supervised fine tuning sft skill.