Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Manage an agent’s context window as a per-request budget: count the complete request in the format your provider receives, reserve room for the response, and remove or compact low-value history before you hit the limit. Keep durable task state outside the rolling conversation so the agent can recover after compaction or a new session.

What a context window limits

A context window is the maximum token capacity available to a single model request—not a quota for how much conversation your application may store. In an agent loop, the request can contain instructions, prior messages, tool definitions and results, retrieved documents, and multimodal inputs, as well as tokens used to generate the response. The exact accounting depends on the model and API: OpenAI includes input and output, and reasoning tokens for some models; Anthropic counts the system prompt, messages, tools, and generated output; Gemini describes a combined input/output limit. Check the current documentation for the model and endpoint you use rather than relying on one universal capacity figure. See OpenAI’s conversation-state guide, Anthropic’s context-window guide, and Google’s token guide.

Context capacity and the endpoint’s output-token limit are related but distinct constraints. A request that leaves too little capacity for generation can produce an incomplete answer even if its input was accepted. Reasoning models may also use part of the context for reasoning. Set the output limit deliberately and handle incomplete responses; do not assume a nearly full prompt leaves enough room to finish. OpenAI explains its reasoning-token accounting in its reasoning-model guide.

Count what the API will receive

Tokens are not words. Tokenization varies with the model, encoding, language, and content type, so word counts and text-only estimates are not reliable measures of a full request. Message boundaries and roles, tool definitions, schemas, images, and files can all affect request size.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the provider’s model-matched counting method on the complete request shape, then compare estimates with the usage reported after calls. OpenAI’s token guide describes token counting and the limits of plain-text estimates; its API documentation also describes counting input for the Responses API. For Gemini, use count_tokens and the model-information interface for the model and request you intend to use, as described in Google’s token documentation. Log actual input and output usage, including cached-token fields when the provider returns them. This gives you a way to spot estimates that drift from real usage and requests that grow unexpectedly.

Set a budget and act before the hard limit

Build a budget around the next request, not just the latest user message. Estimate or count all included context, reserve enough capacity for the expected answer and any applicable reasoning, and set the endpoint’s output limit to fit that plan. There is no evidence-based universal trigger percentage: choose one using observed request sizes, response needs, and how much recovery time an overflow would cost your application.

  1. Measure: Count the full request with the provider’s matching model and API method.
  2. Compare: Check the estimate against the context and output limits currently documented for that model and endpoint.
  3. Act early: When the request approaches your chosen trigger, prune, retrieve, split, summarize, or compact context before sending it.
  4. Validate: After the call, log reported usage and whether the response completed. Adjust your trigger and estimates when production behavior shows they are too optimistic.

Model limits and API behavior can change between models or snapshots. Verify the current model reference during implementation rather than turning a model-specific capacity into a timeless rule.

Keep active context useful as history grows

More context is not automatically better. Keep information that helps the current task and reduce material that merely makes the request larger. Useful actions include:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Remove duplicated instructions, repeated conversation excerpts, and stale tool output.
  • Retrieve or select relevant documents for the current step instead of repeatedly including an entire corpus.
  • Divide oversized source material into manageable parts, then preserve the findings needed for later steps.
  • Summarize older conversation when continuity matters, retaining concrete facts, constraints, decisions, and unresolved questions rather than vague narrative.

Long-context support does not eliminate cost, latency, or retrieval-quality concerns. Google’s long-context guidance notes that performance depends on the workload and discusses caching when large inputs are reused. Whether retrieval, summarization, or a larger context is the right choice depends on what information later steps must reliably access.

Choose compaction with provider behavior in mind

Compaction replaces or compresses prior conversation state so later requests can continue with less context. It is provider-specific; a compacted result is not necessarily a plain-text summary that you can edit or reconstruct. Follow the provider’s continuation rules and test what information survives for your task.

Provider or feature What the documentation describes Implementation consideration
OpenAI Responses API Compaction can be triggered server-side at a configured rendered-token threshold or by a separate compact operation. The returned compaction item is opaque and encrypted. With input-array chaining, append returned items and you may drop items preceding the latest compaction item. With previous_response_id, pass only the new user message; do not also prune history manually. Follow the state-chaining pattern in OpenAI’s compaction guide.
Anthropic API The documentation describes threshold-based compaction using context_management.edits and a beta strategy. Later requests continue from the compaction block while earlier blocks are dropped. Check current beta availability and model coverage, and carry forward the returned block as documented. See Anthropic’s compaction-threshold guide.
OpenAI Agents SDK sessions Sessions persist conversation history. OpenAIResponsesCompactionSession can replace longer stored history with a shorter item list; its documented default trigger is item-count-based and can be customized. Choose token-based or other heuristics if item count is a poor proxy for size. Avoid combining this compaction session with a server-managed conversation session that uses a different history flow. See the Agents SDK sessions guide.

Before adopting any provider’s compaction feature, confirm its current availability, supported models, and continuation semantics. A threshold controls when compaction happens; it does not guarantee that every detail important to your task will be retained.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Persist task state beyond the rolling context

Conversation history is useful working context, but it is a weak sole source of truth for long-running work. Store critical state in a session, database, or explicit artifact so a new process or session can resume without relying on the model to recall every earlier message. OpenAI’s conversation-state guide and Agents SDK sessions guide describe persistence options; Anthropic’s context-window documentation also addresses preserving state across work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A useful state artifact is concise and operational. Include the objective, hard constraints, decisions and reasons, source-of-truth references, completed work, open questions, and the next action. Update it when a meaningful decision changes the plan, and load it into a new session alongside only the context needed for the next step. Keep the authoritative state outside any provider-specific compaction block so you can recover if a session expires or a compaction loses a critical detail.

Monitor for overflow and state loss

Token usage alone does not show whether the agent remains effective. Track usage alongside completion status, latency, and the application’s cost measures. Inspect cases where a response is incomplete or where a post-compaction action omits a required fact. Before continuing a workflow, validate that required state—such as the current objective, constraints, and next step—is present. If it is missing, reload it from the durable artifact rather than asking the model to infer it from a shortened history.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.