Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteTo reduce context usage in a multi-step AI automation, send each model call only the instructions, history, tool definitions, and results it needs for its next decision. Inspect the assembled request first, then trim irrelevant inputs, retrieve large source material on demand, keep tool schemas and outputs lean, and compact stale conversation state when appropriate. Prompt caching can lower the cost of repeated processing, but it does not shrink the context window occupied by the request.
What counts as context in an automation?
Context is the complete model-visible request, not just the text in your prompt. Depending on the application, it can include system and developer instructions, the current user turn, earlier messages, editor or application state, referenced files, tool definitions, and tool outputs. An agent that appears to “send the whole conversation” may also be sending hidden or implicit state assembled by its framework.
That distinction matters: shortening one prompt string may barely affect total usage if large tool schemas, prior results, or attached files remain in every request. Start by inspecting representative requests across several steps, using your provider’s request logs or usage telemetry where available. Attribute input tokens to instructions, history, references, tool definitions, and returned data so you can see what is actually growing.
Reduce what each step receives
Make instructions specific to the task
Use task-specific instructions instead of one universal prompt containing every possible rule. Keep shared instructions for genuinely shared requirements, and supply step-specific guidance only where needed. This makes the request easier to reason about and avoids carrying irrelevant directions through every model call.
#1 Best Overall
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
Pass only relevant files and references
Attach only the files, records, or documents needed for the current decision. For a large corpus, keep source material in a filesystem, database, or retrieval layer and ask the model to open or parse focused portions when needed. OpenAI describes this pattern for an agent working with a computer environment: the model can inspect relevant material through tools instead of receiving an entire environment in the prompt (OpenAI’s Responses API computer environment overview).
The design question is not “How do I compress everything?” but “What must this step see to make the next decision?” Keep exact details available through durable storage or a retrieval pointer when they may be needed later.
Keep tool definitions and results lean
Trim schemas without weakening them
Tool definitions consume context even before a tool is called. Remove redundant wording and expose only the parameters and behavior the model needs, while retaining required fields, constraints, and safety instructions. Test the resulting schema against realistic calls: an overly terse or ambiguous description can cause errors that cost more in retries and corrective turns than it saves.
Load tools on demand where supported
Some platforms can defer tool definitions until they are relevant. Anthropic’s Claude documentation describes tool search as a way to avoid placing a large toolset in baseline context; its guide suggests considering it when a toolset grows past roughly 20 tools or baseline context use becomes noticeable. That is a vendor heuristic, not a universal cutoff. Tool discovery may add a lookup step, so weigh its context savings against latency and call count. See Anthropic’s tool context documentation for platform-specific behavior.
Keep intermediate results out of the transcript when possible
Tool outputs often accumulate in conversation history. Return concise, structured summaries with identifiers or retrieval pointers when later steps can fetch full details as needed. For sequences of small deterministic operations, batch work in the application or use a provider feature that keeps intermediate results outside the conversational transcript. Anthropic documents programmatic tool calling for this purpose, along with context editing that can remove stale tool results. These capabilities and their semantics vary by provider; verify support for the model and API path you deploy.
Do not discard information that a later decision depends on. Prefer retaining a compact result containing the outcome, relevant identifiers, and any essential constraints, while leaving bulky logs or raw records in retrievable storage.
Rank #3
Compact long-running state deliberately
When a conversation has become long, compaction can replace accumulated history with a smaller continuation state. OpenAI documents server-side threshold compaction and a standalone compact endpoint. For the standalone endpoint, use the returned output as the canonical next context; for server-side compaction, follow the documented input-array or response-ID chaining pattern instead of manually pruning the request. Provider formats and support differ, so use the exact continuation method documented for your API.
A useful compacted state preserves the information the next step cannot safely reconstruct:
Free tools Windows power users keep installed
One-click scans. No signup required.
- The objective and active constraints.
- Decisions already made and why they matter.
- Exact identifiers, code snippets, library choices, and completed actions with outcomes.
- Open questions, blockers, and the next action.
Amazon Bedrock’s compaction guidance, for example, calls out preserving code snippets, library choices, and retry and rate-limit decisions. Validate critical values against durable application state rather than relying on a summary for exact data.
Compaction has costs as well as benefits. AWS says Bedrock compaction requires an additional sampling step that contributes to billing and rate limits, and that compaction may be followed by a cache miss. Measure whether the context saved on later requests outweighs that overhead for your workflow. See Amazon Bedrock’s compaction documentation.
Keep reusable prefixes stable, but do not confuse caching with less context
Prompt caching reuses processing for a matching request prefix and may lower repeated input cost; cached tokens still occupy context. Anthropic puts the distinction plainly: “Prompt caching doesn’t reduce the number of tokens in context, but it reduces what you pay for them on subsequent requests.”
OpenAI recommends placing stable developer instructions and shared reference material first, with dynamic content such as timestamps and user-specific values later. Append new turns rather than rewriting old ones when the workflow allows it. A stable leading prefix can improve the chance of a cache match, but it does not guarantee one; rewriting history, compaction, or truncation can change the prefix. See OpenAI’s prompt caching guide.
Best Value
OpenAI’s documentation says cached input tokens may be discounted by up to 95%; the documentation page does not state a publication year, and the actual discount depends on model pricing. Treat that as a possible pricing benefit, not a promise of lower context occupancy or a universal savings rate.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Separate unrelated jobs and pass a focused handoff
Conversation history is often scoped to a session, and context may not carry into a different session automatically. When an automation switches to unrelated work, start a new session rather than bringing along an irrelevant transcript. If work must continue elsewhere, pass a short handoff with the task, constraints, decisions, current result, blockers, and next action.
For application-specific context, Microsoft’s guide explains the kinds of material an agent request may assemble, including instructions, conversation, references, and tool results. Use that as a checklist, then verify the actual payload and usage details for your provider and framework: Understand context in AI agents.
Measure context, compaction, and cache use separately
Track input or context token counts per step, compaction usage or charges where exposed, and cached-input tokens independently. A lower bill may indicate cache reuse rather than a smaller request. OpenAI’s prompt caching guide describes cache diagnostics and notes that discount rates vary by model pricing. Without separate measurements, cost alone cannot show whether model-visible context has fallen.
| Approach | What it changes | Main trade-off |
|---|---|---|
| Selective references and retrieval | Removes irrelevant source material from the current request. | Later steps may need a retrieval call to fetch details. |
| Lean tool definitions and concise results | Reduces baseline schema context and accumulated tool output. | Over-trimming can make calls ambiguous or lose needed detail. |
| Compaction or context editing | Removes or summarizes prior material in the continuing context. | Summaries can lose exact detail; compaction can add work and affect cache continuity. |
| Prompt caching | Reuses processing for a matching prefix; does not reduce context occupancy. | Prefix changes can interrupt reuse, and savings depend on model pricing. |
Choose based on the bottleneck. If the request is too large for the model, prioritize removing or retrieving material and compacting stale state. If repeated input cost is the problem and the context fits, stable-prefix caching may help. If a workflow depends on exact details, preserve them in durable state rather than trusting a lossy summary. OpenAI’s compaction guide, Anthropic’s tool-context guidance, and Bedrock’s documentation describe provider-specific implementation paths; verify current model, region, SDK, and API support before deploying them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

