Glossary · term

Scaling laws

Neural scaling laws are empirical relationships that estimate how a model's held-out prediction loss, typically validation or test cross-entropy, changes as parameters, training data, and compute increase. They describe measured regularities within a specified regime and can support forecasts for larger training runs. The related expression scaling wall is an informal, contested hypothesis that further pretraining scale may face sharply diminishing practical returns or binding resource constraints; it is not part of the technical definition or a demonstrated universal stopping point.

Training2017-12-01Wave 1 · 2023Maturity: 4/5

Origin and context

Hestness and colleagues reported predictable scaling behavior across deep-learning applications in 2017. Kaplan and colleagues then documented power-law relationships across language-model size, dataset size, and training compute in 2020. Hoffmann and colleagues later found that, under a fixed compute budget, many large language models had been trained on too little data and proposed a different compute-optimal balance. These revisions illustrate what scaling laws do: summarize measurements within a regime and guide resource allocation, rather than prescribe one permanent recipe.

Sources: s5, s1, s2

Why it matters

Scaling laws let research teams estimate the likely held-out loss from a proposed training run, compare allocations before spending a large compute budget, and reason about where additional resources may help. The wall question matters because usable data, chips, power, capital, and communication latency may constrain a forecast even when a fitted curve still improves. Data-constrained experiments show that repeated data can help but does not erase data limits, while Epoch AI's analysis treats power, chips, data, and latency as separate bottlenecks rather than evidence of one settled wall.

Sources: s5, s2, s3, s4

Example

A team can train several smaller models, fit a held-out-loss-versus-compute curve, and estimate the data and compute needed for a larger run. If that estimate requires more high-quality tokens or power than can be obtained, the team has encountered a planning constraint. Calling it a scaling wall should remain shorthand for the constraint and uncertainty, not a claim that all model improvement has ended.

Sources: s5, s1, s3, s4

How it differs

Compute Wall / Data Wall

Scaling laws describe measured performance trends and support forecasts. A compute wall or data wall names particular resource bottlenecks that can prevent a forecasted run from being practical. A resource wall may therefore bind even when the empirical loss curve has not flattened.

Maturity and evidence

The empirical scaling-law concept is established across independent research groups and supports maturity 4, although its coefficients and compute-optimal prescriptions change with methods, data, and evidence. The separate scaling-wall label remains contested and underspecified; its inclusion as a related debate does not lower the maturity assigned to scaling laws themselves.

Sources: s5, s1, s2, s3, s4

Limits and open questions

A scaling law fitted to held-out cross-entropy loss does not guarantee a corresponding gain on every downstream capability, and extrapolation outside the measured range can fail. Dataset composition, architecture, optimization, post-training, and evaluation choices can move the curve. Evidence for a power law in one regime neither proves indefinite progress nor proves a universal wall.

Sources: s5, s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as model training skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as distributed training skill.