context window
A context window is the span of tokens an AI model can use within a single request, including input and generated output.
A context window is the span of tokens an AI model can use in one request. It covers the input and the output being generated. A prompt is the input or instruction given to the model; the context window sets a limit on how much information the model can handle at once.
This entry is based on AIPOST articles and widely known facts. If something is wrong, please send us a correction request.
Articles covering this entry
The returned policy lands in the model's context window, the working memory it draws on for the current task, so the code it writes next follows company require…
As a session runs, the context window, the amount of text a model can consider at once, fills up with conversation, instructions and tool output.
Training treats the context window, the text the model can see at once, as if it were the state of the world.
…on November 19, 2025, with GPT-5.1-Codex-Max, a model trained to work across multiple context windows, the limit on how much text a model can consider at once.
The model's context window, the amount of information it can consider at once, is 128,000 tokens.
…Cache Clock mod puts that information in the status bar: how much of the context window, the amount of conversation the model can hold at once, is in use, and h…
Their tool descriptions are added to every prompt, taking up space in the context window, the amount of text a model can consider at once.
The API model listing also gives GPT-6.1 Sol a 1,050,000-token context window, a 128,000-token output limit and an April 30, 2026 knowledge cutoff; it accepts t…
…t describes page structure and controls, but a complex page can consume substantial space in the model’s context window—the information it can consider at once.
Luna supports text and image inputs, has a 1,050,000-token context window, and can produce up to 128,000 output tokens.
Codex also compacted its context about 20 times per active hour; compaction condenses earlier material so work can continue within a limited context window.
Repeated corrections also fill the agent’s context window—the limited amount of information it can use at once—with conversational history.
A local agent used Qwen3.8 27B with a context window of 131,131 tokens (the amount of material the model can consider at once) and quickly answered questions dr…