Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

The reported token savings in one August 2026 benchmark range from 91% to 93.9% depending on what is counted. Those percentages describe different comparisons: one contrasts full conversation history with assembled memory context; the other compares billed input on a hand-picked 27-question subset. Neither is an answer-quality result or a production guarantee.

What the three token totals count

A September 28, 2026 post on DEV Community, published under the account “belcore,” reports three figures from the same benchmark work. They are not interchangeable: two count context, while the third counts billed input for a limited subset.

Reported figure What it counts Scope
109,079 tokens per call The full, uncut prior transcript Context count, using the o200k_base tokenizer
6,671 tokens per call The assembled memory context Context count, using the o200k_base tokenizer
9,908 billed input tokens per call Input including the prompt and question A hand-picked subset of 27 questions; the post describes this result as indicative only

The first two figures produce the post’s 93.9% context-token reduction. The roughly 91% billed-input reduction uses a different basis: the 27-question subset and input that includes the prompt and question. Calling both figures “token savings” without naming their scope makes the comparison easy to misread.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the benchmark was set up

The author says the dataset was LongMemEval_S, containing 500 questions and about 109,000 tokens of prior conversation per question—roughly 500 turns. The answer model was GPT-5, the harness was frozen, and a fixed seed was used. Context counts were tokenized with o200k_base. The runs took place in August 2026.

These details define the author’s reported setup; they do not establish that the figures will transfer to another model, tokenizer, conversation mix, or application. The post also describes the benchmark as public and not production traffic.

What the percentages establish—and what they do not

The reported reductions describe token counts under this setup. They do not establish that answers remained equally accurate, because the post does not report answer quality for this comparison and says accuracy is a separate metric. Nor does it establish the total cost per successful task: retries and fixes were not included in a measured end-to-end cost.

The roughly 91% billed-input figure deserves particular care because it comes from 27 questions selected by the author, not a random sample of all 500. The post labels it indicative only. A reader evaluating a similar system would need to know the count’s scope, whether prompt and question tokens are included, how questions were sampled, the answer-quality result, latency, and total cost per successful task including retries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the multi-pass retrieval experiment adds

The post also describes a secondary experiment: retrieval made another pass to check for and retrieve evidence before answering. Results differed substantially by sample.

  • Hand-picked missing-evidence subset: correct answers increased from 16 to 20 among 27 questions.
  • Random sample: the attributable effect was +2 among 103 questions, within a ±5.2 noise bar. In one case, extra evidence changed a correct aggregation answer into a wrong one.
  • Resource use on the subset: the author reports about 2.9 times the input tokens and 2.6 times the latency.

The stronger result on the hand-picked subset should not be treated as the expected result on ordinary questions. The author says this approach is not being shipped as an improvement.

How much confidence to place in the accuracy observations

The author reports that 38% of audited wrong answers involved a gold label judged wrong or defensible either way. This is an observation from the author’s audit, not independent verification. The post also notes that n = 499 cannot resolve small effects. Together with the limited samples and the absent answer-quality result for the token-saving comparison, these caveats constrain what can be concluded from the percentages.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare token-saving claims fairly

When evaluating a claim, match the measurement to the question you actually care about:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Token scope: Is the count the full transcript, selected context, or billed model input?
  • Included material: Are the system prompt, other prompts, and the current question counted?
  • Sampling: Were questions randomly sampled or deliberately selected for a particular case?
  • Answer quality: Were answers checked, and was accuracy measured on the same questions?
  • End-to-end cost: Does the figure include retries and fixes needed to complete a task successfully?
  • Latency: How much time does the approach add or save?

For the figures in this post, context-token counts and the subset’s billed-input count are reported, along with selected multi-pass accuracy and resource-use observations. An answer-quality result for the main token-savings comparison and a cost-per-successful-task measurement including retries and fixes are not reported.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.