Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

If the total on your AI usage dashboard does not match your provider invoice, the software doing the arithmetic is a likely suspect. An audit described in an article by Roy Tong, published on Dev.to on 22 September 2026, examined 110 open-source tools that count tokens, track costs, or enforce budgets. The article reports more than 45 verified bugs across those tools, grouped into five recurring error families, and 23 fixes merged upstream. These are the audit authors’ claims as of their September 2026 snapshot. They have not been independently reproduced in this article.

The short answer to why a usage bill is wrong: the meter inherits one of five faults. It may use a stale price table, apply a cache multiplier meant for another provider, count retries twice or deduplicate real work away, record missing usage as zero, or break its quota logic at a calendar boundary. Each fault can produce a total that looks precise and is still wrong.

What the audit covered and what it does not establish

The article says the work took about a month and covered open-source projects that count tokens, track costs, or enforce budgets. Read its figures with four limits in mind:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The results are the audit authors’ claims, dated to their September 2026 snapshot.
  • Sample-based figures, such as the seven-tool pricing sample and the 20-vendor survey, describe those samples. They are not measures of how common each problem is across the ecosystem.
  • A secondary summary by AI Vibe News, dated 23 September 2026, describes code reading, synthetic fixtures and parser runs. That methodology description has not been checked against the authors’ own report.
  • The full list of 110 tools, the linked issues and the pinned commits were not available for this article. Tool-specific findings should therefore be read as the authors’ claims until checked against their report and code.

The five error families

The article groups its bugs into five families. Each one can survive a code review and still produce a number that looks plausible.

Stale pricing tables

A usage meter multiplies token counts by per-model rates taken from a pricing table. If that table lacks a model you now use, or still carries a rate you no longer pay, every downstream total inherits the error, and the output gives no warning. In the article’s sample of seven tools, five had outdated or missing pricing rows.

Cache multipliers applied to the wrong provider

Cached prompt tokens are often billed differently from ordinary input, and cache reads and cache writes can carry different treatment. The article says these values vary by provider, so a multiplier that is correct for one vendor can be wrong for another. Its example is an Anthropic cache-read discount applied to OpenAI models, which understated cache reads by five times. That is one configuration error, not a ratio that holds across providers, and it says nothing about current list prices.

Retries, streams and double counting

Yes, a token tracker can double-count retries. When a request is retried, the stream can re-emit events identical to ones already counted, and naive aggregation adds them again. The article reports 604 re-emitted events in a public corpus it cites, of which 46% were byte-identical. The opposite mistake is also possible: deduplicating too aggressively can erase real work. A sound meter separates logical operations, meaning what the application asked for, from physical attempts, meaning what reached the provider. It also applies a deduplication rule that is explicit and can be audited.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why a dashboard shows zero when usage happened

Missing usage is not zero usage. Some code paths turn an absent usage field into a numeric zero, so a request whose response omitted token counts appears free. The article’s proposed correction is to keep absence as absence, so a rollup can report that a total is unprovable rather than inventing a zero. If a day with real traffic shows zero cost or zero tokens, check whether the underlying records lacked usage fields before trusting the figure.

Quota windows that fail at calendar boundaries

Quota windows are often anchored to wall-clock time, such as a daily or monthly reset. Tests that pin fixtures to absolute dates can pass when written and fail later as the calendar moves past those dates. The article treats this as a boundary-condition family and recommends testing quota windows relative to time boundaries. Its description of this family is general, and this article does not attribute a specific reproduced defect to it.

How to check whether a cost meter is accurate

The table below turns the five families into checks you can run against a meter’s exported records or its source code.

Check What a sound meter does Symptom when it fails
Pricing version Records the model and pricing-table version behind each cost figure A cost cannot be reproduced after a rate change
Cache handling Applies cache-read and cache-write rules keyed by provider and model Cached-token costs differ from the provider’s billing for the same model
Retries Separates logical operations, retries and physical attempts, with a documented deduplication rule Totals rise after retries, or events disappear during aggregation
Missing fields Keeps absent usage as unknown and flags it in rollups Days with traffic show zero cost or zero tokens
Quota windows Tests windows relative to time boundaries Tests pass today and fail after the date rolls over
  1. Export one day’s usage records for a period when you know the traffic occurred.
  2. Recalculate the cost for that day by hand, using the provider’s published rates for the same models and dates.
  3. Check each record for a missing usage field, a retried request and a cache-read count, and note whether the meter handled each one as the table describes.
  4. Investigate any gap between your recalculated figure and the meter’s total before using the meter for budgets or billing reconciliation.

What the local conformance pack does and does not prove

The article describes an open-source conformance pack released under the MIT license. Its stated workflow is to export usage data and run the checker locally, so no data leaves the machine. The checker separates logical operations from physical attempts, cache reads from cache writes, and absent values from zero. It ties each verdict to a named rule. The article also refers to a settlement specification called AMS-1. These are the article’s descriptions. The article does not offer an independent comparative evaluation of competing tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The conformance suite contains 236 checks, and the article says an independent auditor reproduced it. The limits matter. The pack examines exported records and meter logic. It cannot verify provider-side data that was never exported, it does not prove that a provider’s invoice is correct, and it does not establish that every tool is covered.

If you are comparing meters, the article points to these evaluation axes:

Rank #4
Plaud Note Pro AI Voice Recorder Transcribe & Summarize for Meetings Calls
  • ENHANCED CONTEXT WITH MULTIMODAL INPUT: Capture audio, type notes, add images, and press to highlight key moments for richer context. During recording, instantly mark key moments with a single button press. Simultaneously enrich your audio by snapping photos of important documents or typing in ideas
  • CHAT WITH YOUR RECORDINGS USING "ASK Plaud": Unlock deeper insights with this interactive AI. Ask questions, extract key points, draft emails, and get next-step suggestions—all grounded in your original audio for reliable, ready-to-use answers
  • INTELLIGENT RECORDING WITH AI DIRECTIONAL AUDIO: Enjoy seamless, intelligent recording with Plaud Note Pro. Its AI automatically switches between call and meeting modes while recording, while directional audio and real-time spatial awareness minimize noise to capture voices with crystal clarity
  • Everything Included: Includes Plaud Note Pro, magnetic case, magnetic ring, charging cable, and a free Starter Plan with 300 transcription minutes per month. Upgrade anytime in the Plaud app to Pro Plan (1,200 min/mo) or Unlimited Plan(Up to 24 hours of transcription per user per day)
  • PREMIUM ULTRA-SLIM DESIGN WITH INSTANTVIEW DISPLAY: Meticulously designed, the AI Note Taker is just 0.12 inches thin and 1.06 oz —about the size of a credit card. Its sleek aluminum body with a textured wave finish features a vivid AMOLED display, letting you check battery and recording status at a glance, while it seamlessly works with Apple Find My to ensure you never misplace it
  • Pricing maintenance by provider and model
  • Cache read and cache write handling
  • Retry and deduplication semantics
  • Treatment of missing fields
  • Quota-window boundary behavior
  • Local versus remote data handling
  • Traceability from verdicts to rules
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What vendors have published about billing errors

The article reports a cache-accounting problem found by an independent auditor, as relayed by the article’s authors. On an affected commercial-provider cache-accounting path, usage was under-reported by up to 98.9%. The auditor and the provider are not named in the accessible reproduction, so this figure cannot be attached to a specific company here.

The authors also surveyed 20 commercial vendors and found that none had a published dispute or correction process. That is a finding about public process. It does not show that no vendor has a private escalation route, so it should not be read as evidence about how individual billing complaints are handled.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The authors propose two baseline controls: a named dispute path and machine-checkable billing disclosures. They present these as recommendations, not as a published standard or legal requirement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.