Context rot
Context rot is a descriptive label for reduced or less reliable language-model performance as the supplied input becomes longer, even when the input remains within the advertised context window. The effect can vary with task, model, relevant-information position, distractors, semantic similarity, and document structure. It is an observed evaluation pattern, not a diagnosis of one internal mechanism and not a claim that every longer prompt is worse.
Origin and context
Chroma's July 2025 report used context rot for results across controlled retrieval, conversational memory, and repeated-word tasks, varying input length while attempting to hold task difficulty constant. The label builds on an earlier evidence base. Lost in the Middle showed that models could use information at the beginning or end of a long input more reliably than information in the middle. RULER found that nominal context capacity could exceed effective performance on more demanding long-context tasks. Anthropic independently used and defined context rot in April 2026 guidance about managing Claude Code sessions and long contexts.
Why it matters
A large context-window specification tells developers how much input a model accepts, not how reliably it will use every part of that input. Retrieval systems, document assistants, and long-running agents can therefore remain under the hard token limit while still losing accuracy as irrelevant evidence, competing passages, or accumulated history grows. Teams need task-specific curves across length and structure, and should treat the effective context budget as an empirical property of a system rather than a vendor number.
Example
A question-answering service succeeds when one relevant paragraph appears in a short prompt but degrades after many plausible distractors are added. That is evidence consistent with context rot if the team holds the question and target evidence constant and repeats the test across lengths and positions. A single failure caused by an ambiguous question is not enough. The service can compare reranking, truncation, retrieval, and compaction against the same evaluation set.
How it differs
Long-context language models
Long context describes the capacity or engineering of models that accept large inputs. Context rot describes performance degradation observed as inputs grow. A model can accept a long sequence without using it uniformly or reliably, so nominal window size and effective context are different measurements.
Context compaction
Compaction intentionally reduces accumulated context by summarizing, collapsing, or removing material. It can mitigate context pressure, but a lossy summary can create a separate failure. Context rot is the measured degradation pattern; compaction is one context-management response, not its definition.
Maturity and evidence
Maturity is rated 3. The named report is recent, but it synthesizes a phenomenon supported by independent peer-reviewed work and a separate benchmark, while Anthropic has independently adopted the same label in operational guidance. The rating stays below 4 because context rot has no standard metric or causal theory, evaluations cover limited tasks and model snapshots, and usage of the label remains broader than any one experimental setup.
Limits and open questions
Length often changes task difficulty, topic mixture, and distractor count at the same time, making causal attribution difficult. Synthetic retrieval tests can overestimate useful long-context reasoning, while one benchmark threshold cannot define every application's effective window. Model updates can also change results quickly. Claims should name the tested model, task, prompt construction, lengths, positions, and metric rather than convert context rot into a universal percentage or fixed cutoff.
Related terms
References
- Context Rot: How Increasing Input Tokens Impacts LLM PerformanceChroma · 2025-07-14 · class A
- Lost in the Middle: How Language Models Use Long ContextsTACL / ACL Anthology · 2024-02 · class A
- RULER: What's the Real Context Size of Your Long-Context Language Models?NVIDIA / COLM / arXiv · 2024-04-09 · class A
- Using Claude Code: session management and 1M contextAnthropic / Claude · 2026-04-15 · class A
Last updated: 2026-09-04
This term is also covered in the Skills Atlas as long context modeling skill.
This term is also covered in the Skills Atlas as context engineering skill.