An AI agent score is a measurement claim: it says that a system achieved a particular result on a defined task, under a defined scoring method. To make that claim meaningful, specify what the score is meant to measure and how it was produced before running the evaluation. Then check whether the result reflects the intended capability—or whether the agent could access solutions or exploit a gap in the grader.
What does an agent score actually measure?
Start by naming the capability or outcome the evaluation is intended to represent. “Agent performance” is too broad on its own: readers need to know whether the target is, for example, completing a specified task, following constraints, or producing a result judged against explicit criteria. The task set and scoring method should make that target concrete.
This is a question of validity, not just arithmetic. A score can be calculated consistently and still fail to measure the capability its label implies. NIST describes a related failure mode as occurring “when an AI model exploits a gap between what an evaluation task is intended to measure and its implementation, solving the task in a way that subverts the validity of the measurement.” NIST’s CAISI page on cheating in AI agent evaluations discusses this risk alongside contamination, where an agent can access solutions to evaluated tasks.
Benchmark quality also affects interpretation. BetterBench assesses AI benchmarks, underscoring that a benchmark name or a precisely computed number is not, by itself, evidence that the evaluation supports the intended conclusion. There is no single metric or threshold established as suitable for every agent evaluation; the choice depends on the evaluation’s objective and method.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
What should be fixed before the evaluation runs?
Write down the protocol before looking at results. That reduces the chance of silently changing what counts as success after seeing which systems perform well. A useful protocol records:
- Target and task scope: the capability or outcome being measured, the tasks included, and what counts as a valid completion.
- Metric and scoring procedure: how the result is computed, what criteria are applied, and who or what assigns the score. If a grader is used, describe its role and the rubric it follows.
- System configuration: the model or agent tested, any scaffolding, and the tools or other affordances it can use.
- Restrictions and conditions: what the agent may not do, plus the environment and run protocol in which it is evaluated.
- Interpretation boundary: what a result supports—and what it does not establish about performance outside this benchmark and setup.
Scoring criteria should be inspectable enough for a reader to understand how a broad goal becomes a result. OpenAI’s PaperBench offers one example: it uses hierarchical rubrics divided into individually gradable subtasks. OpenAI reports 8,316 such tasks and says paper authors helped develop the rubrics. It also reports a separate benchmark for evaluating its LLM judge. Those are methodological choices in PaperBench, not proof that any rubric or language-model judge is valid by default.
Rank #2
How can a high score mislead?
Solution contamination
If evaluated solutions are accessible to an agent, success may reflect access to those solutions rather than the capability the benchmark is intended to assess. NIST identifies contamination as a risk to evaluation validity. A report should therefore state what contamination controls were used and avoid implying that a score proves independent problem-solving if solution access has not been ruled out.
Grader loopholes
A scoring implementation can reward behavior that satisfies its checks without accomplishing the intended task. This is different from contamination: the problem is a gap between the task’s goal and the grader’s criteria. Inspect the rubric and scoring logic for that gap, and report what was checked. A reproducible result under a stated scoring procedure is useful, but reproducibility alone does not show that the procedure measures a real-world outcome.
Rank #3
How should two published scores be compared?
Compare the evaluation objective as well as the evaluation process. The ACM survey on evaluation and benchmarking of LLM agents distinguishes these dimensions; NIST’s discussion also highlights the importance of agent affordances and restrictions. Use the following checklist before treating two results as comparable:
- Do they target the same capability and cover similar tasks?
- Are they using the same benchmark and dataset version?
- Are the tested model or agent and its scaffolding comparable?
- Did the systems have the same tools, affordances, and restrictions?
- Were the run protocol and environment materially alike?
- Do the metric, rubric, and grader measure success in the same way?
- Are contamination controls described, and is uncertainty or repeat-run treatment reported?
If important setup or scoring details are undisclosed, treat the scores as results from different evaluations rather than a clean ranking. A shared benchmark label cannot resolve differences in task scope, tooling, or grading.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What does a benchmark-specific result look like?
OpenAI reported that its best-performing tested system on PaperBench was Claude 3.5 Sonnet (New) with open-source scaffolding, with an average replication score of 21.0% on that benchmark. The figure describes that tested system, scoring setup, and PaperBench task; it is not a general measure of agent capability or a prediction of performance on other tasks. See OpenAI’s PaperBench report for the benchmark context.
When publishing any score, keep the claim at the same scope as the evidence: identify the benchmark and tested configuration, explain how the score was produced, and distinguish performance under that protocol from evidence about deployment outcomes.
Recommended Free Tools
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

