Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
An AI agent is working when it reliably completes the task it was assigned, at a quality and latency you can accept, for a cost that makes sense for each completed outcome. A polished final answer does not prove that. To tell the difference, read the full execution trace, grade the result against criteria you wrote down in advance, and track cost per successful task alongside success rate and latency. Vendor documentation supports this approach, but it does not supply a universal ROI formula or a pass threshold, so the thresholds below are choices you will need to make for your own workload.
Define what “working” means before you measure anything
Most teams start with cost, because the invoice is the visible problem. Cost only means something once you know what the agent was supposed to deliver. Write each task type as an observable outcome. Useful questions include: Was the requested action completed? Did the result pass a domain-specific check? Did the agent call the required tool? Did it respect the instructions it was given?
Build a representative set of tasks rather than judging a few memorable runs. OpenAI’s “Evaluate agent workflows” documentation describes a progression: inspect traces to find failure modes, write structured graders for them, then assemble a dataset so you can rerun the same evaluation after every change. That sequence matters because a prompt tweak that improves one vivid example can quietly degrade the other ninety percent of traffic.
Do not rely on a single model-judge score as the final verdict. Combine structured graders or deterministic checks (for example, “the ticket status field now reads closed”) with human review of a sample. The documentation supports graders and evaluation runs, but it does not prescribe one correct mix, so the right balance depends on how costly a wrong answer is.
#1 Best Overall
Read the full run, not the final answer
A trace records the steps inside a turn: model responses, tool calls, delegated work, step status, duration, and the data recorded at each step. OpenAI’s tracing documentation describes this structure, and it is the fastest place to see where an agent went wrong or where it burned time. A confident final sentence can sit on top of a failed tool call that the agent quietly worked around, or of an unfinished task it described as complete.
When you open a representative trace, look for these patterns:
- Failed or erroring tool calls that were retried, bypassed, or ignored.
- Repeated or unproductive calls, such as the same search issued several times with slightly different wording.
- Incorrect tool choice, where a cheap lookup was skipped in favour of an expensive or unnecessary step.
- Unnecessary handoffs to another agent when a single agent could have finished the work.
- Long steps whose duration dwarfs the rest of the run.
Trace grading can score these workflow-level questions directly: whether the agent chose the right tool, handed off when appropriate, followed instructions, or improved after a prompt or routing change. OpenAI’s evaluation guidance puts it plainly: “Trace grading is the fastest way to identify workflow-level issues.”
Rank #2
Count the whole run’s cost, not the visible answer
Estimating spend from the length of the final response will undercount almost every agent run. OpenAI’s observability and usage documentation identifies the cost-bearing components below. Check each one against your own accounting.
| Cost component | What it covers | Common miss |
|---|---|---|
| Input tokens | Everything sent to the model on each call | Context that grows across a long multi-step run |
| Cached input tokens | Input served from cache, which is priced differently from fresh input | Treating all input as one rate |
| Output tokens | Text the model generates | Counting only the visible answer |
| Reasoning tokens | Hidden reasoning the model uses; OpenAI bills these as output | Not present in the visible response, so easy to overlook |
| Multiple model calls | Every call made over the course of one task | Summing only the first call |
| Root agent and subagent work | Delegated work performed by other agents in the same task | Measuring only the top-level agent |
| Retries | Calls repeated after errors or unsatisfactory output | Logging only the final successful call |
| Tool, sandbox, and third-party charges | Non-model costs incurred during the run | Not stated in model-usage dashboards by default |
The practical rule is simple: attribute every cost to the task run that caused it, including the failed attempts. A run that costs little but needed three retries is not cheap.
Make cost per successful task your headline number
Cost per run tells you what the agent spends. Cost per successful task tells you what it spends to deliver value. Calculate it over a representative window:
Rank #3
total attributable run costs ÷ number of tasks that passed the agreed success criteria
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →This is an editorial metric recommendation drawn from the vendor guidance above, not a published industry standard. Always report success rate and latency next to it. Otherwise a cheaper configuration that fails more often looks efficient when it is not.
Here is an illustrative calculation with hypothetical figures, to show why the denominator matters. Both configurations handle 200 tasks in the same window:
Rank #4
| Configuration | Total attributable cost | Tasks passing criteria | Success rate | Cost per passing task |
|---|---|---|---|---|
| A (larger model, fewer retries) | $60.00 | 150 | 75% | $0.40 |
| B (smaller model, more retries) | $30.00 | 120 | 60% | $0.25 |
Configuration B looks cheaper per attempt and still looks cheaper per passing task, but it leaves 30 tasks unfinished, and those failed tasks still cost money. Whether B is the better choice depends on what a failed task costs your business: a support escalation or a rework step may cost far more than the $0.15 difference per success. Your numbers will differ; the calculation is what carries over.
Check whether your instrumentation can support the number
A cost dashboard is an estimate until you reconcile it. Its accuracy depends on complete usage data and correct model and price metadata. Two widely used tools handle this differently.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Langfuse
Langfuse’s token and cost tracking documentation describes two routes. You can ingest usage and cost reported in the model response, or you can infer cost from a model definition and the recorded usage. Inference has a hard limit: for reasoning models such as OpenAI o1, Langfuse states that it cannot determine the correct cost when token usage is missing, because it cannot see the reasoning-token count. For those generations, you must supply token usage explicitly.
Best Value
LangSmith
LangSmith’s cost tracking documentation describes automatic cost calculation from token counts and model prices for supported LLM calls. For other run types, including tools and retrieval, you enter costs manually. For a custom calculation, your application must supply the token counts, the model or provider information, and the price. If any of those three is missing, the cost figure is incomplete.
Provider usage data
OpenAI’s tracing documentation notes that recorded usage counts can arrive after a turn has ended, and that a blank or null value means unknown, not zero. Recorded values may also change. Those counts are not necessarily the final bill. Before you use operational telemetry as a financial record, reconcile it against provider billing data for the same period.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Compare configurations on the same dataset
When you change a prompt, model, routing rule, or tool configuration, run the old and new versions against the same evaluation dataset with the same criteria. Otherwise you are comparing different workloads. For each configuration, record:
- Pass rate against the defined success criteria.
- Cost per passing task, with the cost components and estimation method stated.
- Latency, including which steps in the trace are slowest.
- Failure, retry, and tool-call patterns.
How much spend a task deserves depends on its value. A high-stakes task, such as a compliance check or a customer-facing refund, may justify more calls and higher cost for reliability. A low-value internal summary may not. Set the acceptable cost per successful task per workflow, not once for the whole system.
Warning signs that the agent is burning money
- The answer reads well, but the trace shows a failed action or an unfinished step. Grade the task outcome, not the text.
- Calls and retries are climbing while success stays flat.
- Cost totals look suspiciously low because tool, retry, delegated, or non-model charges are missing from the accounting.
- A monitoring tool infers model cost without the token usage it needs, especially for reasoning models.
- A dashboard shows blank usage and the team reads it as zero.
When the numbers disagree: a short troubleshooting path
- Cost per passing task rose but success rate held. Look for added retries, longer traces, or extra handoffs that did not change outcomes. Trim the steps that do not change the result.
- Success dropped after a change. Compare failing traces before and after the change, focusing on tool choice and instruction-following. Roll back if the regression is clear.
- Cost looks too low to be true. Check whether subagent work, retries, and non-model charges are recorded, and whether usage fields are null for some runs.
- Reasoning-model costs are missing. Confirm that token usage, including reasoning tokens, is being ingested rather than inferred from an incomplete record.
No published, cross-industry statistic on agent success rates or wasted spend was identified in the official documentation reviewed, so treat any figure you see elsewhere with caution. Your own trace-based baseline is the number that matters.
The sources behind this article are OpenAI’s “Evaluate agent workflows,” “Observability and usage,” and “Tracing” documentation, Langfuse’s “Token & Cost Tracking” documentation, and LangChain’s “Cost tracking” guide for LangSmith. These describe vendor capabilities and implementation behavior. They are not independent comparative tests, and this article does not assert current prices or plan limits for any product.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

