Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteStructure an AI agent’s context in layers: keep durable goals and rules in its instructions, pass the current task and relevant conversation state into the model-visible history, maintain selected long-term facts as memory, and retrieve changing or extensive knowledge only when it is needed. Keep application state separate unless the model needs to see it, and treat retrieved content as untrusted input.
How do I structure context for an AI agent?
Start with a distinction that is easy to miss: an application may hold data that the model cannot see. OpenAI Agents SDK documentation puts it this way: “When an LLM is called, the only data it can see is from the conversation history.” Instructions, tool results, retrieved passages, and other information become available to the model only when the application surfaces them in the interaction.
Context engineering is therefore broader than writing a good prompt. Anthropic describes it as curating the full set of tokens available to a model during inference. In an agent loop, that set changes as the agent receives tool results and new messages, so context needs to be selected and refreshed across turns.
| Context source | Best role | Design consideration |
|---|---|---|
| Instructions | Stable goals, behavioral policy, constraints, and output requirements | Keep transient facts and large reference collections out of this layer. |
| Application and runtime state | Dependencies, authorization, identifiers, and current structured state | It is not automatically model-visible. Surface only the fields needed for the task. |
| Conversation input and history | The immediate user request and relevant recent turns | Long histories can become costly or distracting; summarize or prune when appropriate. |
| Persistent memory | Selected user preferences, durable learnings, and compact notes | Maintain it, check freshness, and resolve conflicting updates. |
| Retrieval and tools | Large, changing, or on-demand external knowledge and actions | Check relevance and provenance, and handle returned content as untrusted. |
Keep instructions stable
Use instructions for information that should guide many tasks: the agent’s purpose, enduring constraints, behavioral policies, and required response format. Avoid turning instructions into a dump of documents, current records, or user-specific facts. Those items change at different rates and are usually better supplied through runtime state, memory, or retrieval.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Separate application state from what the model sees
Your application may know which account is active, what permissions apply, which records are available, or what step of a workflow is in progress. The model does not inherit that knowledge merely because it exists in application memory. Pass only the relevant, authorized details into the model-visible context, and keep access-control decisions enforced by the application rather than relying on the model to infer them.
Make the current task and necessary history visible
Include the user’s immediate request and the portions of prior conversation needed to answer it. If older turns no longer matter, omit them; if decisions or facts must carry forward, summarize them without silently changing their meaning. The useful context is not necessarily the entire transcript, but the state needed to handle the present step correctly.
Retrieve external knowledge when it is useful
Use tools or retrieval-augmented generation (RAG) for information that is too large, too changeable, or too task-specific to keep in the instructions or every request. The retrieved evidence must still be added to the model-visible context if the model is expected to use it. More context is not automatically better: irrelevant, conflicting, or excessive material can make a response worse.
Rank #2
What should go in an agent’s memory versus its prompt?
Use the prompt or current conversation context for information needed now, and persistent memory for a small set of facts likely to remain useful across future tasks. “Prompt” here means the model-visible input for a call, not only the user’s latest sentence. It can include instructions, relevant history, selected application state, and evidence returned by retrieval.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Put it in persistent memory when… | Put it in the current context when… |
|---|---|
| It is a durable preference or learning likely to help across sessions. | It is specific to this request, conversation, or workflow step. |
| It can be stored compactly and maintained as facts or notes. | It is current state that should be verified for this particular run. |
| It is still useful after the current exchange ends. | It comes from a changing source and should be fetched when needed. |
Memory is not a second, unquestioned source of truth. Preferences and notes can become stale, and new information can conflict with old entries. Prefer current verified state when it disagrees with an older memory; update or retire the note rather than accumulating contradictory versions.
There is no single standardized memory format. OpenAI Agents SDK documentation describes a pattern of extracting conversation summaries and raw memory notes, then consolidating them into a more usable layout. AWS Prescriptive Guidance recommends combining structured state and recent dialogue with summaries and retrieval from long-term memory. These are implementation patterns, not requirements for every agent.
How should retrieval work for an agent?
For a knowledge base too large or changeable to include directly, retrieval selects evidence for the current question and adds it to the model’s context. In Anthropic’s 2024 description, a common RAG pipeline divides a corpus into chunks, creates embeddings for semantic similarity search, and supplies relevant chunks to the prompt. Each part addresses a different need: chunking makes material searchable in pieces, while embeddings help find conceptually related passages.
Semantic search may not reliably surface an exact product code, name, or identifier. Anthropic notes that lexical matching such as BM25 can help catch exact phrases and identifiers that embeddings may miss. Combining lexical and semantic retrieval, deduplicating results, and reranking candidates are options described in its approach—not universal requirements. Choose them based on the corpus and the queries the agent must handle.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Anthropic reported that its Contextual Retrieval method reduced failed retrievals by 49%, and by 67% when reranking was added. Those figures describe Anthropic’s reported method in its 2024 publication; they are not independent benchmark results or a promised improvement for other systems.
Choose retrieval for the workload
Compare retrieval and direct inclusion against the actual demands of the task, rather than adopting a token threshold as a rule. Anthropic’s 2024 article says direct inclusion may be simplest for some knowledge bases below 200,000 tokens in the Claude context discussed there. That is a model- and publication-specific example, not a universal cutoff. Google’s Gemini API guidance cautions that long-context performance can vary when a task requires finding multiple targets, and that longer inputs can increase latency and cost.
- Corpus size and update frequency: Frequently changing or extensive material is a strong reason to retrieve on demand.
- Exact-match needs: Identifiers and exact strings may require lexical matching alongside semantic search.
- Relevance and retrieval quality: Measure whether the system returns the evidence needed, not merely whether it returns something.
- Latency and token cost: Retrieval adds work; including large amounts of text also consumes context and can affect latency and cost.
- Data sensitivity: Limit what gets retrieved and shown to the model to the information the task is authorized to use.
- Consequence of error: Higher-impact actions call for stricter evidence checks and review.
How do you assemble context for each run?
Build the model-visible context from the current task outward. The following sequence is a practical pattern, not a required SDK-specific message format.
- Load stable instructions. Supply the agent’s durable role, policies, constraints, and output requirements.
- Resolve runtime state in the application. Determine the active workflow state, authorization, relevant identifiers, and dependencies. Do not assume the model can see these values.
- Form the current task context. Add the user’s request and only the conversation history or maintained summary needed to interpret it.
- Retrieve what is missing. Use tools or search for relevant external evidence, then include useful results with enough source information to assess them.
- Check before acting. Validate tool arguments and permissions in application logic; route consequential operations for review when appropriate.
- Update state deliberately. After the run, retain only memory that is likely to be useful later, and keep summaries and stored facts consistent with newer verified information.
For repeated static context, caching may be available in some products, but support and economics depend on the current model and pricing terms. Check the relevant current documentation before designing around a caching feature or assuming it will reduce cost.
Best Value
How should you test context quality and agent reliability?
Evaluate retrieval and the model’s response as separate stages. OpenAI API accuracy guidance identifies two distinct failure modes: retrieval may supply missing, noisy, or excessive context, and the model may still misuse relevant context. If an answer is wrong, determine whether the needed evidence was absent or poor, or whether the model failed despite receiving good evidence.
- Test retrieval: Does the system find relevant, sufficiently complete evidence for representative questions? Does it handle exact identifiers and changing information?
- Test response use: Given good evidence, does the model answer accurately, respect constraints, and avoid unsupported claims?
- Test context size: Does adding more history or retrieved text help, or does it introduce noise, latency, or unnecessary token use?
- Test memory maintenance: Does the agent use durable facts appropriately, and does current verified state take precedence when a note is stale?
- Test consequential workflows: Check behavior with realistic authorization boundaries and tool outputs, including hostile or misleading content.
How should an agent handle untrusted retrieved content?
Retrieved pages, files, and tool outputs can contain instructions intended to manipulate the agent. OpenAI security guidance notes that prompt injection may arrive through web pages, retrieved files, or MCP and file-search outputs, and that model defenses do not catch every attack. Treat such material as evidence to assess, not as authority to override system policy or application controls.
Reduce risk in layers: use trusted integrations and carefully chosen file sources, constrain available capabilities, validate tool arguments with schemas or other checks, and log or review tool calls where warranted. In workflows where public research and sensitive data access intersect, separating those capabilities can reduce exposure. For consequential actions, require appropriate human review or application-side authorization rather than trusting a model’s interpretation alone. These measures reduce risk; they do not guarantee that an agent will detect every attack.
What is the right amount of context?
There is no universal prompt layout, memory schema, or retrieval threshold that fits every agent. Choose a design for the workload, then test it with representative tasks. Long-context capacity does not remove tradeoffs: quality, retrieval accuracy, latency, and cost depend on the model and task. A compact, relevant context that is refreshed as state changes is often more useful than indiscriminately supplying everything the application knows.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

