compaction
Compaction is a technique that summarizes and compresses earlier conversation context so a language model can keep working past its context window limit.
Compaction is a technique used by agents built on large language models during long tasks. When the conversation and work history approach the limit of the context window, the earlier material is summarized or compressed so that only the essential state remains, and the compressed version becomes the starting point of a fresh context window.
It is used in long-running coding agent work to preserve earlier requirements and progress while reducing token usage. A known limitation is that details can be lost during compression.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
Harnesses handle this with compaction, which condenses a long history into a compact summary.
What makes these long runs possible is compaction, which keeps the essential project state and drops the noise as a thread grows.
Cache Clock + Compact: the cache timer and compaction button described above
Codex also compacted its context about 20 times per active hour; compaction condenses earlier material so work can continue within a limited context window.