Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A coding agent’s token bill is the sum of every model request it makes during a task, and each request can re-send a growing body of context. Repeated instructions, tool definitions, conversation history, file contents, and tool results all count as input on every turn. Tool-call arguments and reasoning count as output. A Reddit user’s comparison of a simple two-file edit reported about 760 KB of JSON exchanged by one agent workflow and about 100 KB by a single-shot workflow. Those are byte counts of serialized payloads, though, not tokens and not dollars. To learn where your own usage goes, you need the provider’s usage records and a trace of each request.

What the Reddit comparison actually measured

The post, by a Reddit user identified as cgouguen, describes a deliberately simple PyQt task in a two-file project: change the card width to the total width divided by three. The author ran it through an agent workflow called Pi and through Aider, which handled the edit in a single request. The post’s date could not be independently confirmed, so treat it as an undated account.

Measure (as reported by the author) Pi agent workflow Aider single-shot edit
Model requests (LLM calls) 3 1
JSON exchanged, measured by the author About 760 KB About 100 KB
What the model received A request to read both files, the full file contents returned by the harness, then results from later edit steps A preassembled prompt containing a repository map, the raw text of both files, formatting instructions, and the user’s request
What the model returned Several edit tool calls, then a summary of the result One answer containing SEARCH/REPLACE blocks
Provider-reported input and output tokens Not stated Not stated
Invoice or usage export Not stated Not stated

The author stresses that the task was unusually simple and that the comparison fits cases where the developer already knows which files to edit. The payload sizes come from one task, measured by one person, and were not reproduced in a controlled benchmark. They show how much material moved between the tool and the model. They do not show how many tokens were billed, or in what ratio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The personal bill claim

The author also says a monthly API bill above $400 fell to under $100 after part of the workflow moved to single-shot edits for known files. The post offers no invoice, token export, controlled workload, or breakdown that isolates how much of the drop came from that change. It is a before-and-after account from one user, not a predictable saving.

Where the tokens come from

An agent’s cost accumulates across model requests, not only in the final visible answer. OpenAI’s usage documentation lists the input sources that can be billed: agent instructions, tool definitions, conversation history, user input, files or images, and tool results. Output includes the visible answer, tool-call arguments, and reasoning. OpenAI’s documentation states: “Reasoning tokens are billed as output tokens.” A trace can therefore show a large input burden from repeated or growing context even when the visible answer is short.

Input: context you pay for again on every turn

Each model request processes everything the harness sends with it. In a multi-step run, earlier tool results are carried forward, so the input for the third request usually contains the material from the first two plus whatever came back in between. Anthropic’s pricing documentation likewise describes tool definitions and returned tool results as additional consumption, and notes that tool versions can carry different overhead. Those details are specific to each provider and API surface, so check them against the model you actually use.

Output: what the model generates

Output is more than the text you read. It includes tool-call arguments, which can be long when an edit is written as a structured payload, and reasoning tokens where the model exposes them. OpenAI’s documentation says reported output usage includes all generated tokens, including formatting and tool-call structure that may not appear in the message content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tool execution is not automatically a model charge

Running a file read or a shell command does not itself bill model tokens. The problem is that the returned text becomes input to the next model request. Sandbox, infrastructure, or third-party charges may also apply, and OpenAI recommends including them when estimating what a task costs.

How a multi-step run adds up

The Pi sequence in the post follows a pattern common to many agent loops. Each step below is a separate model request, and each one can carry everything before it.

  1. The first request includes the system instructions, the tool definitions, and the user’s task.
  2. The model asks to read both files. The harness returns their full contents.
  3. The next request includes the task, the tool definitions, and both files, plus the model’s read calls.
  4. The model emits one or more edit calls. The harness returns confirmations for each.
  5. A final generation summarizes the work, and that request carries all of the accumulated history.

Counted this way, the cost of a task is not the size of the answer. It is the sum of the inputs and outputs of every request in the chain.

Reading a trace without guessing

OpenAI’s tracing guide organizes a session into turns and groups the work into agent, generation, and tool spans. Agent spans identify root or subagent work and show usage recorded for that agent. Generation spans hold one model request’s inputs and outputs. Tool spans show each call and its result. Use them in this order:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Group the spans by turn, so you can see which requests belong to the same user task.
  2. Check the agent spans to separate the root agent from any subagents. A root-only view omits delegated work.
  3. Open each generation span and compare its input size against the previous one. Growth from step to step usually points to carried-forward tool results.
  4. Match each tool span to the result that was returned to the model, so you can see which outputs were fed back in.
  5. Read the session usage summary last. It may be delayed, may be unknown, and can change after the turn. A blank or null value means unknown, not zero, and recorded usage is not necessarily the final bill.

Why JSON size and billed tokens diverge

Several things separate a byte count from a billed token count:

  • Serialization. JSON adds syntax, escaping, and structure. Tokenization is model-specific, so the same content can produce different token counts across models.
  • Generated structure. Reported output usage can include formatting and tool-call structure that never appears in the visible message.
  • Reasoning. Reasoning tokens count toward output usage even though they are not shown as ordinary text.
  • Caching. Matching prompt prefixes may be served from cache, but eligibility, prefix matching, and cache lifetime rules apply, and no hit is guaranteed. Cached input is billed at its applicable rate, which can differ from the uncached rate.

Caching deserves a caution of its own. A high cached-input percentage does not prove a lower total cost, because a repeated large history may still be processed on every request. Judge the total usage and the bill, not the percentage.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to measure a fair comparison

Comparing two workflows requires the same task, the same model, the same configuration, and the same quality threshold. Then record the following for each request, using whatever the provider exposes:

Measure Why it matters
Model identifier, run or session, turn, agent or subagent, step Lets you attribute each request to a point in the workflow
Input tokens and cached input tokens, reported separately Shows how much context was processed and how much came from cache
Output tokens and reasoning tokens, where exposed Separates visible answers from generated structure and reasoning
Price schedule and billed usage Connects token counts to the amount you actually pay
Retries and delegated agent work Prevents undercounting work that happened outside the main path
Tool, sandbox, and observability charges, listed separately Keeps model usage distinct from other costs in the total
Task success under the same quality threshold A cheaper run that misses the change is not a saving

Be explicit about your accounting boundary. The OpenAI Agents SDK tracks usage for each API request and aggregates it across calls within a run. Persistent sessions can feed earlier messages back in as input on later runs, so a per-run total and a session-level total answer different questions. Label which one you are reporting before comparing workflows.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a single-shot edit is the right trade-off

  • Known files and a narrow change. If you already know which files and locations to edit, a single request can skip exploratory reads and extra turns. This is the situation the Reddit author describes.
  • Discovery and iteration. An agent still earns its cost when it must find the relevant code, run tests, react to failures, or revise its plan. Forcing a single request onto that work can produce a cheap run that fails.
  • Prompt stability. Keeping instructions and tool definitions stable can improve the chance of cache reuse. Confirm the reuse in the usage data rather than assuming it.

For a narrow edit on known files, a single-shot workflow plausibly reduces the number of requests and the amount of re-sent context. How much it saves on your account depends on your model, your prompts, and your quality bar. The Reddit figures do not establish that result, and a trace of your own runs is the only way to confirm it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.