Glossary · term

Continual pre-training (CPT)

Continual pre-training (CPT) resumes a language model's self-supervised pre-training on one or more later corpora instead of rebuilding the model from the beginning. The aim is to absorb new domains, languages, or time periods while retaining useful earlier capabilities. CPT names a training process, not one algorithm: replay, learning-rate schedules, regularization, and parameter-isolation methods can all be part of it.

Training2019-07-29Wave 2 · 2024Maturity: 3/5

Origin and context

ERNIE 2.0 used continual pre-training in 2019 for incrementally learning pre-training tasks. ALUM used the wording in 2020 for adversarial training while continuing from an already trained language model. Gong and colleagues applied the exact phrase to domain adaptation for mathematical problem understanding in 2022; Ke and colleagues later used Continual PostTraining for sequences of unlabeled domains. In 2023, independent teams studied continual domain-adaptive pre-training and learning-rate re-warming. These papers document related but not identical recipes.

Sources: s5, s6, s4, s3, s1, s2

Why it matters

A model may need newer knowledge or better coverage of a domain after its original training run. Continual pre-training can reuse the existing checkpoint and direct compute toward the new corpus. The engineering problem is not merely resuming a job: a distribution shift can improve performance on new data while degrading performance on earlier data. Teams therefore need replay or other retention measures, explicit data lineage, and evaluations spanning both the incoming and original distributions.

Sources: s6, s1, s2, s3

Example

Suppose a general language model must learn a new collection of scientific papers. A team can continue the pre-training objective on that collection, mix in selected earlier data, re-warm and then decay the learning rate, and test both scientific tasks and a regression suite for general capabilities. Training only a small supervised adapter for one downstream label set would instead be fine-tuning, even if both projects start from the same checkpoint.

Sources: s6, s1, s2

How it differs

Post-training

Post-training is a broader and inconsistently bounded phase that can include instruction tuning, preference optimization, or reinforcement learning after broad pre-training. Continual pre-training specifically continues a pre-training-style objective on later corpora. Some papers use post-training for this operation, so reports should name the objective and data rather than rely on the label alone.

LoRA and QLoRA

LoRA and QLoRA update low-rank adapters while keeping most base weights fixed. Continual pre-training describes when and why training continues, and it may update all weights or use parameter-efficient components. The concepts can be combined, but neither is a synonym for the other.

Maturity and evidence

Maturity is rated 3. Independent research teams have published concrete objectives, schedules, replay strategies, and evaluations, and the problem has persisted across several model and corpus settings. The term is not standardized, however, and evidence remains sensitive to model scale, distribution shift, retained-data access, and the meaning assigned to CPT.

Sources: s5, s6, s1, s2, s3

Limits and open questions

Published gains do not establish that one re-warming or replay recipe transfers to every model or corpus shift. Earlier data may be unavailable for replay, and aggregate benchmarks can hide forgetting in narrow capabilities. Continual pre-training changes model weights rather than attaching an external knowledge source, so teams should compare new-domain gains with regressions on retained distributions and preserve the earlier checkpoint for rollback.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as continual pre training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as training data curation skill.