iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Context engineering is the practice of deciding what information an AI agent receives at each step—and what it should retain, summarize, discard, or fetch later. It goes beyond prompt wording: instructions, tool definitions, conversation history, tool results, retrieved evidence, and prior outputs can all shape the model’s next decision. The practical goal is not to fill the context window, but to supply the most useful information for the task at hand.
What is context engineering?
Context engineering is the ongoing design and maintenance of the information available to a model during inference. Anthropic describes it as curating the tokens that land in the model’s context, including information beyond the prompt itself: Effective context engineering for AI agents.
Prompt engineering usually focuses on the wording and structure of instructions. Context engineering includes that work, but also asks which other information belongs in the active context, how it gets there, and when it should be removed or preserved elsewhere. For an agent that takes many steps, this is an iterative process: each tool call and result can change what the model needs next.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsWhat counts toward an agent’s context?
A context window is the information a model can reference while generating a response, including the response itself. Its contents are broader than the visible user conversation. Anthropic’s context-window documentation says inputs, outputs, tool configuration, and, where applicable, thinking tokens count toward the limit.
#1 Best Overall
- Instructions: system-level guidance, task requirements, and constraints.
- Tools: tool definitions and descriptions that explain available actions.
- Conversation history: user requests, assistant messages, and prior decisions.
- Tool results: file contents, search results, API responses, and other returned data.
- Retrieved evidence: external material selected to inform the current step.
- Generated output: the model’s response also uses part of the available window.
That accounting matters for tool-heavy agents: a short user-visible exchange can still create a large context if the agent reads files, searches broadly, or receives verbose API responses. Window capacities and accounting details vary by model and provider, so check the target model’s current official documentation rather than treating one provider’s limits as universal.
Does a larger context window make an agent better?
Not automatically. A larger window can make more information available, but capacity is not the same as relevance. Anthropic characterizes context as a finite resource with diminishing returns and notes that accuracy and recall can decline as token count grows. Irrelevant or repetitive material can compete with the evidence the model needs for its next decision.
Relevance and placement matter alongside size. Keep stable instructions and constraints available, but retrieve task-specific evidence when it becomes useful rather than loading a whole knowledge base by default. A practical target is the smallest context that gives the agent enough information to make the next decision reliably.
How do I manage context for a long-running agent?
Start by identifying what the agent needs at each inference step: its goal, non-negotiable constraints, relevant decisions so far, available tools, and evidence needed for the next action. Then choose a context-management technique based on what is causing the context to grow or what information must survive.
Use retrieval for large or changing knowledge
Retrieval brings selected external information into the active context when it is needed. It is useful when the agent needs access to a large corpus but only a small portion is relevant to a given decision. Anthropic describes embedding-based retrieval as a common pre-inference approach and discusses just-in-time context strategies in its agent context engineering guidance.
Retrieval is not a substitute for state management: it can find source material, but the agent still needs a concise account of its current task and decisions. Retrieve focused evidence and preserve where it came from when that provenance matters.
Use compaction when conversation history is too long
Compaction summarizes a long interaction so the agent can continue with a shorter, high-fidelity representation. A useful summary preserves the goal, constraints, key decisions, unresolved questions, and implementation details needed to resume. It should not silently turn uncertain assumptions into established facts.
Free tools Windows power users keep installed
One-click scans. No signup required.
Compaction addresses history growth; it does not necessarily solve an oversized tool result or the need to carry knowledge across separate sessions. Anthropic’s agent cookbook discusses context-management patterns including compaction.
Rank #3
Clear old tool results that can be fetched again
If raw file reads, search results, or API responses dominate the context, remove old outputs that can be retrieved again. Keep the useful conclusion or a reference to the source, rather than repeatedly carrying the full result forward. This can reduce context bloat without discarding the fact that a tool was used.
Clearing is a poor fit for information that cannot be recovered or for evidence the next step depends on. Anthropic’s cookbook describes context editing and tool-result clearing patterns; the right policy depends on the agent’s actual tools and workload.
Use external memory for knowledge that must persist
Persistent memory stores selected information outside the active context so it can be retrieved later or carried across sessions. It is suited to durable project facts, preferences, or progress notes—not indiscriminate transcripts. In Anthropic’s described memory-tool approach, the application developer controls the storage backend.
Memory introduces decisions about what to save, how to retrieve it, and how to update or remove stale information. It is different from compaction: compaction shortens the current interaction, while external memory preserves selected knowledge beyond that active context.
Use structured notes and focused subagents for long-horizon work
For work likely to span resets or sessions, structured notes can record current status, completed work, open questions, and next actions. These notes help reconstruct state without replaying every prior message. Focused subagents can handle bounded subtasks with a narrower context, then return results to the main agent.
Both methods depend on good task decomposition and a clear handoff. A note that omits a key constraint or a subagent that returns conclusions without relevant evidence can make the overall system less reliable.
Which context technique should you try first?
| Observed problem | Technique to try | What it changes |
|---|---|---|
| The conversation history is crowding out current work. | Compaction | Replaces a long interaction with a focused summary of goals, decisions, and open work. |
| Large, repeatable tool outputs dominate the context. | Tool-result clearing | Removes old results that can be fetched again while retaining necessary conclusions or references. |
| The agent needs a small amount of information from a large corpus. | Selective retrieval | Fetches relevant evidence when needed instead of loading the corpus by default. |
| Important information must survive a session or context reset. | External memory or structured notes | Stores selected knowledge outside the active context for later recovery. |
| A large task can be split into bounded, independent parts. | Focused subagents | Gives each subtask a narrower context and returns a handoff to the main agent. |
These techniques can be combined, but add them to address diagnosed problems rather than by default. For example, an agent may retrieve source material selectively, compact its history when needed, and store only durable project state externally.
Recommended Free Tools
How should you evaluate context changes?
Test changes against the same representative workload and tool-use pattern. First identify whether the bottleneck is history length, verbose tool output, access to external evidence, or continuity across sessions. Then compare the original configuration with one targeted change.
Best Value
- Task performance: Does the agent complete the intended work correctly?
- Reliability: Does it preserve constraints and handle errors consistently?
- Token use: Does the change reduce unnecessary context without removing necessary evidence?
- Latency: Does retrieval, summarization, or memory lookup add meaningful delay?
- State quality: Can the agent resume accurately after compaction or a reset?
- Engineering and control: What storage, retrieval, or update logic must your application maintain?
Anthropic’s cookbook recommends diagnosing which part of context growth is responsible and testing clearing configurations against the workload’s tool-use pattern. A change that saves tokens in one workflow may remove useful information in another.
What do Anthropic’s evaluation figures show?
Anthropic reported results from internal evaluations in a 2025 announcement about context editing and memory. On its internal agentic-search evaluation, it reported a 29% improvement over baseline for context editing alone and a 39% improvement when it combined its memory tool with context editing. In a separate 100-turn web-search evaluation, Anthropic reported an 84% reduction in token consumption using context editing. These are vendor-reported results for Anthropic’s evaluations, not independent replications or guarantees of similar gains in other systems. See Anthropic’s context-management announcement for the reported results and evaluation context.
A broader research survey also indicates the field’s scope, not a direct comparison of techniques: the authors of the 2025 preprint A Survey of Context Engineering for Large Language Models say they reviewed more than 1,400 research papers. That paper count is the authors’ description of their review, not an independently verified census.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

