Glossary · term

Context compaction

Context compaction reduces the active history sent to a language model while preserving enough state for a conversation or agent task to continue. A system may summarize older turns, collapse bulky tool results, remove low-value history, or replace earlier context with a compact state object. Compaction manages an inference-time context budget; it does not enlarge the model's native context window or update model weights.

LLMOps2025-03-18Wave 3 · 2025–26Maturity: 3/5

Origin and context

Anthropic's Claude Code changelog associates automatic conversation compaction with version 0.2.47, which npm registry metadata dates to March 2025. Anthropic later described compaction as one response to context pollution in long-running agents, alongside structured note-taking and multi-agent architectures. OpenAI documented native Responses API compaction in March 2026, using a compacted item plus selected recent context. Microsoft's Agent Framework separately documented truncation, sliding-window, tool-result, and summarization strategies. The shared label covers several representations and policies rather than one interoperable format.

Sources: s4, s6, s1, s2, s3

Why it matters

Agent loops accumulate user turns, tool calls, outputs, plans, and intermediate evidence. Sending all of it can exceed a hard window and can also raise token cost, latency, and the amount of irrelevant material the model must navigate. Compaction makes long-running work operationally possible by choosing what survives. That choice is consequential: a summary that omits a constraint, unresolved decision, citation, or tool outcome can silently change later behavior.

Sources: s1, s2, s3

Example

A coding agent approaches its context threshold after many searches and test runs. Its compactor preserves the user's goal, file boundaries, accepted decisions, current failures, and a concise record of tool outcomes, while removing superseded logs. The next model call receives that compact state and recent turns. The team then tests whether constraints and pending work survive repeated compactions, not only whether the prompt became shorter.

Sources: s1, s2, s3

How it differs

Context Engineering

Context engineering is the broader discipline of selecting, structuring, securing, and maintaining all information supplied at inference time. Compaction is one technique within it, normally applied after history accumulates. A context design can use retrieval, memory, or delegation without compacting a transcript.

Active Context Curation

Active context curation continually decides what evidence should enter or remain in working context. Compaction specifically reduces accumulated state under a budget. Curation may drive compaction, but it can also add newly retrieved material or replace stale evidence rather than summarize history.

Prompt Caching

Prompt caching reuses computation for an unchanged prefix. Compaction changes the representation or selection of context so fewer tokens remain active. One reduces repeated compute; the other reduces or restructures content, and a system may use both.

Maturity and evidence

Maturity is rated 3. Multiple independent platforms and an open framework expose concrete compaction mechanisms, making the term operational rather than hypothetical. It remains below 4 because interfaces are recent, meanings differ across systems, and there is no shared measure of fidelity or standard for which state must survive.

Sources: s1, s2, s3

Limits and open questions

Every compaction policy is lossy unless it retains a fully reversible representation. Summaries may erase provenance, exact wording, negative results, security boundaries, or dependencies that later become important. Repeated summarization can compound omissions. Opaque platform-native items can also reduce portability and auditability. Teams should preserve external durable state, test adversarial and long-horizon cases, and keep hard constraints outside disposable narrative history when possible.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-04

In the Skills Atlas

This term is also covered in the Skills Atlas as context engineering skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as long context modeling skill.