iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
You stop wasting tokens on a long-context model by deciding, request by request, what the model actually needs to see. Put stable material where a provider cache can reuse it, fetch only the passages a question depends on, and compress or summarize history only after a task-level test shows the answer still holds. Caching lowers the price of a repeated prefix but does not remove the cost of new tokens, and retrieval, compression and summarization each add their own failure modes. Savings depend on the model, the workload and the implementation, so the number that matters is the one you measure on your own traffic.
What context engineering covers
Context engineering is the work of selecting, arranging, transforming and maintaining everything that enters an LLM’s input, across a single request or a long session. Prompt wording is only one part of it. A 2025 survey, A Survey of Context Engineering for Large Language Models, organizes the field around retrieval and generation, processing, and management. Its authors report that their systematic analysis covers more than 1,400 papers. The survey treats retrieval-augmented generation (RAG), memory, tool-integrated reasoning and multi-agent systems as broader implementations of the same discipline. The survey is on arXiv.
That reframes the practical question. “How do I write a shorter prompt?” is less useful than “Which tokens earn their place in this request, and what happens to them as the session grows?”
Why a longer window does not mean better context
Extended inputs cost more than their token count suggests. They enlarge the key-value (KV) cache, the per-token memory the model holds while it processes a sequence, and they place more material in front of attention, where relevant passages compete with irrelevant ones. Long-running agents make this worse. Old tool output, superseded decisions and stale notes accumulate, so the model works with a context that is large and partly wrong for the step in front of it. A larger window does not make retrieval or memory management unnecessary, because long context still has relevance and retrieval limits.
#1 Best Overall
Five levers, five different mechanisms
“Reducing tokens” is not one technique. Each method below changes a different part of the pipeline, fails in a different way, and needs its own test. Caching changes what you pay for while the model sees the same input. The other four change what the model sees, which is where quality risk enters.
| Method | What it changes | Best question to test | Main risk |
|---|---|---|---|
| Prompt or context caching | Reuses prior computation for a matching prefix or cached content | Do many requests share stable content, and does the cache actually hit? | Prefix drift, prefixes that are ineligible or too short, provider-specific limits, cache misses |
| Retrieval (RAG) | Selects a subset of external information for each task | Does the selected context preserve answer quality at lower total cost? | Missing relevant evidence, retrieval overhead, multiple model calls |
| Prompt compression or token dropping | Shortens the representation supplied to the model | Does the compressed prompt keep the detail the task depends on? | Loss or distortion of key facts |
| Compaction and structured memory | Summarizes or carries forward state across a long interaction | Can the next phase continue correctly from the retained notes? | Omitted decisions, stale summaries, a changed cache prefix |
| Larger context window | Allows more input in a single request | Does full-context access improve the target task enough to justify its cost? | More irrelevant content, memory and cost load, long-context retrieval failures |
Measure before you change anything
- Fix the model and the inputs. Use the exact model you will deploy and a sample of real requests. Cache minimums and cache behavior vary by model and request settings, so figures from one model do not transfer to another.
- Break each prompt into its parts: stable instructions, reference documents, tool output, conversation history and the new user input. Count the tokens for each part. A single total hides where the waste is.
- Record a baseline: input tokens, output tokens, any cache write and cache read counts your provider reports, latency, and a task-specific quality score on a fixed evaluation set. Score against the task itself, such as correct answers, required citations or passing tests.
- Change one variable at a time, whether that is caching, retrieval depth, compression or compaction, and rerun the same evaluation set.
- Compare total cost and quality against the baseline. Prompt length alone is not the metric (see the section on re-measuring the whole system).
Prompt caching: what it does and when it pays
What caching actually changes
OpenAI’s prompt-caching documentation states the core idea directly: “Prompt caching reuses work when requests share the same prompt prefix.” The prompt is still sent, and the new part of the request still has to be processed. Caching reuses model-side state for an eligible repeated prefix, so the saving applies to that prefix and not to the request as a whole. Caching is not token deletion.
The conditions for a cache hit
- The rendered prefix matches. A change in wording, ordering or serialization before the breakpoint means the cache does not match.
- The prefix ends at an eligible cache breakpoint.
- The cacheable prefix meets the model’s minimum length. For GPT-5.6 and later, the minimum cacheable prompt is 1,024 tokens. For earlier models the minimum varies with request settings. Hidden system tokens do not count toward the minimum.
- The model and request settings are the same across the requests you want to match.
The pricing arithmetic
For GPT-5.6 and later, OpenAI’s page lists cache writes at 1.25 times the standard uncached input rate and cache reads at 0.1 times for most listed models. For GPT-6.1 Sol it lists a cache read at 0.05 times. These are one provider’s relative rates for specific models. Check the live pricing for the model you call before you budget.
Rank #2
| Event | Multiplier on the standard input rate | Cost of a 20,000-token prefix, in standard-rate token equivalents |
|---|---|---|
| Uncached request | 1× | 20,000 |
| Cache write (first request) | 1.25× | 25,000 |
| Cache read, most listed GPT-5.6-and-later models | 0.1× | 2,000 |
| Cache read, GPT-6.1 Sol | 0.05× | 1,000 |
This is a hypothetical calculation, not a measured result. Suppose 100 requests share the same 20,000-token prefix and every request after the first hits the cache. Without caching, the prefix costs 2,000,000 token equivalents. On most listed models with caching it costs 25,000 + (99 × 2,000) = 223,000, about 89% less. On GPT-6.1 Sol it costs 25,000 + (99 × 1,000) = 124,000, about 94% less. These figures cover only the cached prefix. New suffix tokens, output tokens and any extra requests are billed separately.
The same multipliers show the break-even point. One write followed by one read costs 1.35 times the standard rate for that prefix, against 2 times for two uncached requests, so a single reuse recovers the write premium. That holds only if the second request actually matches the first.
Order the prefix so it can be reused
- Place stable system instructions and rules first.
- Add reference material that changes rarely, such as a policy text or a product manual.
- Put request-specific content last: the user’s question, the current record, and fresh tool output.
- Serialize the same content the same way on every call, with identical whitespace, key order and formatting.
- Do not rewrite earlier content in place. Editing an old message changes the prefix, so requests after the edit cannot reuse the earlier cached prefix.
Choosing between caching, retrieval and a longer window
These options answer different questions. Start with the shape of the workload.
Rank #3
- The same large material serves many requests unchanged. Test caching first. Google’s long-context guide for Gemini calls context caching “the primary optimization when working with long context and the Gemini models,” and describes caching uploaded files for repeated chat-with-your-data requests. That guidance is specific to Gemini. Its prices and cache behavior do not carry over to other providers.
- Each request needs a different slice of a large corpus. Test retrieval. Compare top-k or chunk selection against a fuller-context baseline on your own questions. Google notes that retrieval accuracy can vary when a request has multiple information targets, and that retrieval accuracy and cost interact. Retrieval is not automatically cheaper or more accurate than full context.
- The answer depends on relationships spread across the whole input. Test the larger window as a baseline. Keep it only if the quality gain justifies the extra input cost. Fitting in the window is not a reason to use it.
These options can be combined, for example a cached instruction block with retrieved passages. Measure the combination as its own configuration instead of assuming the savings of each part will add up.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Compression and token dropping
Compression shortens the representation the model receives. It can work, but it can also remove information or introduce errors, and a shorter prompt does not show that the answer survived. Yuan et al.’s benchmark, “KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches” (Findings of EMNLP 2024), evaluated more than ten approaches from several efficiency families across seven categories of long-context tasks. The authors open with the observation that “no existing work has comprehensively benchmarked these methods in a reasonably aligned environment.” That was the motivation for their 2024 paper.
The practical lesson is to test each method on your own task. Ask a specific question: does the compressed prompt still contain the facts, constraints and exceptions the answer depends on? Check the cases where one missing detail changes the result, such as an exception clause or a number without its unit. An average score can hide those failures.
Rank #4
Compaction and structured notes for long sessions
Anthropic’s engineering article on context engineering for AI agents defines compaction this way: “Compaction is the practice of taking a conversation nearing the context window limit, summarizing its contents, and reinitiating a new context window with the summary.” It also describes structured note-taking, which keeps state outside the live context. Both suit long-horizon work, and both can lose information.
OpenAI’s prompt-caching documentation adds a cache-specific point. Compaction replaces earlier conversation content with a shorter representation, and that can reduce reuse of a prior cache prefix. Its guidance is to compare total input cost before and after compaction, because a lower token count can still save money even when the cache-hit rate falls.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Anthropic’s example of keeping critical details while dropping redundant tool output shows one implementation. It is not a guarantee that a summary is lossless.
Best Value
What a compaction should keep
- Decisions already made, with the reason for each.
- Open questions and the next action they block.
- Constraints and requirements, stated in the user’s terms.
- Essential facts, identifiers, numbers and their units.
- Nothing that repeats a tool log or can be recovered on demand. Redundant logs can be dropped when it is safe to do so.
How to validate a compaction
- Keep the full transcript outside the live context, so a dropped detail can be recovered.
- Generate the summary from a fixed template with the sections above.
- Start a fresh context with the summary only, and ask questions whose answers depend on details that summaries often drop.
- Score those answers against the same questions answered in the uncompacted session.
- Measure total input cost, including the summarization call itself.
Re-measure the whole system
A method’s savings have to be netted against everything it adds. Include these costs in the total:
- Retrieval calls, including any search or embedding work, and the latency they add.
- Summarization or compression calls and the tokens they consume.
- Cache writes, which cost more than standard input, and the cache reads that follow.
- Extra model requests, such as a second pass that checks retrieved context.
- Output tokens and latency, which can change when the context changes.
Fewer prompt tokens do not guarantee lower cost or latency. Suppose a summarization call removes 5,000 input tokens from the main request but costs more than that in its own input and output, and it runs on every turn. The pipeline has then become more expensive, even though the main request is shorter. Compare the full pipeline with the baseline on the same evaluation set.
Quick Recap
Troubleshooting common failures
| Symptom | Likely cause | What to check |
|---|---|---|
| Cache reads are zero or near zero | The prefix changed between requests, or it is below the model’s minimum length | Compare the rendered prefix across two requests and confirm its length against the minimum for your model and settings |
| Some requests hit the cache and others do not | A breakpoint sits after variable content, or request settings differ | Move request-specific content after the breakpoint and hold settings constant |
| Costs rose after adding retrieval | Extra retrieval and generation calls, or a top-k value larger than the task needs | Log the number of model calls per user request and compare top-k values against the fuller-context baseline |
| The answer omits a fact that was in the original input | Summarization, compression or chunking dropped it | Check the missing fact against the full transcript, then add it to the summary’s fact list or retrieve it on demand |
| Cost fell but quality dropped | Compression or compaction removed task-critical detail | Rerun the task evaluation on the uncompressed context and identify which details are missing |
What the current evidence does and does not settle
- There is no general savings figure. No independent number for the tokens or money saved by context engineering as a whole has been established. Plan with your own measurements.
- The rates and minimums above belong to one provider. The OpenAI figures come from its prompt-caching documentation as accessed in 2026. Google’s long-context page was last updated 6 October 2026 (UTC) and covers Gemini only. Neither sets a rule for other providers.
- Some work is a proposal. Teresa Zhang’s “Algorithms for Context Engineering in LLM Inference: Optimization of Placement, Compression, and Scheduling” was published in AAAI proceedings on 14 March 2026. It argues that memory capacity and bandwidth are increasingly limiting, and it frames placement, compression and scheduling as coupled optimization problems. It is an abstract that proposes a framework and a planned evaluation. It does not report measured gains.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

