The Bitter Lesson
The Bitter Lesson is Rich Sutton's 2019 historical thesis that, over long periods of AI research, general methods able to exploit increasing computation—especially search and learning—have tended to overtake approaches built around fixed human domain knowledge. It is an argument about research strategy drawn from selected episodes in AI history, not a theorem, scaling law, or guarantee that more compute wins in every task or time horizon.
Origin and context
Sutton illustrated the thesis with computer chess and Go, speech recognition, and computer vision. In his account, handcrafted domain structure often helped first, but later systems used scalable search or learning to surpass it as computation became cheaper. Sergey Levine invoked the lesson as a persistent theme while discussing scalable learning from large, diverse data, but also noted that reducing it to a slogan can caricature the underlying choices. The title remains tied to Sutton's essay rather than a formal scientific result.
Why it matters
The lesson is used to challenge research plans whose gains depend on ever-growing manual rules, labels, or task-specific engineering. It asks whether a method can continue improving when more compute, data, or search is available and whether human effort becomes the bottleneck. Used carefully, it is a comparative question about scaling paths. Used carelessly, it becomes a slogan that dismisses domain knowledge, safety constraints, data quality, efficiency, or near-term requirements without evidence.
Example
A team can compare two approaches to a perception task: one adds a growing catalogue of hand-written cases, while another learns representations from broad data and improves with larger training runs. The Bitter Lesson favors investigating the second trajectory over the long run. It does not say the first approach has no present value, that data is free, or that the learned system will meet safety and product constraints automatically.
How it differs
Scaling laws
Neural scaling laws are fitted empirical relationships for specified losses, model families, data, and compute regimes; a scaling wall names limits or diminishing returns. The Bitter Lesson is a broader historical interpretation about which research methods benefit from growing resources. Neither logically proves the other.
Test-time compute
Test-time compute gives a model more inference-time search or reasoning work for a request. It can exemplify a general method exploiting computation, but the Bitter Lesson also discusses training and historical search systems. One successful inference technique cannot validate the thesis universally.
Maturity and evidence
Maturity is rated 3. The essay is a stable, attributable reference point and has been taken up in archival research discussions beyond its original page. The score does not rate the thesis as proven: its scope, examples, and practical interpretation remain debatable, and evidence for one scalable regime cannot establish a law across all AI problems.
Limits and open questions
Historical examples are selected retrospectively, and the boundary between general learning and human-designed structure is rarely clean. Compute, data, objectives, architecture, and engineering co-evolve, so a historical comparison cannot isolate one cause. The Chinchilla study shows that resource allocation matters even within a fixed compute budget. The lesson is most useful as a hypothesis to test against alternatives, not as permission to skip ablations or treat scaling choices as self-justifying.
Related terms
References
- The Bitter LessonRich Sutton / UT Austin archival mirror · 2019-03-13 · class A
- Understanding the World Through ActionConference on Robot Learning / PMLR · 2022-01-11 · class A
- Training Compute-Optimal Large Language ModelsDeepMind / arXiv · 2022-03-29 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as model training skill.
This term is also covered in the Skills Atlas as test time compute scaling skill.