Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI benchmark scores are not “total BS,” but they are easy to overread. A score is evidence about a particular task, model setup, and scoring method—not a universal measure of intelligence or a guarantee of performance in your work. The available evidence does not establish that OpenAI or Anthropic generally design benchmarks to trick users. It does show why readers should inspect the conditions behind a score before trusting a ranking.

Are AI benchmarks reliable?

They can be useful when the test matches the claim being made and the evaluation is well designed. A benchmark might measure whether a model answers a set of multiple-choice questions, follows instructions under specified conditions, or completes a particular agent task. Its result supports conclusions about that tested setup; it does not automatically establish how the model will perform in an open-ended conversation or a real workflow.

OpenAI’s guidance on third-party evaluations says strong claims depend on both a suitable harness—the prompts, tools, interfaces, control logic, and other support around a model—and checks that the result is valid. The report should identify what it tested, which system and task distribution were used, and how budgets and elicitation were set. OpenAI’s evaluation playbook also lists contamination, broken problems, reward hacking, refusals, and possible evaluation awareness or sandbagging among risks that can affect interpretation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Evaluation also involves choices about what counts as success. Anthropic’s 2023 comment to the US National Telecommunications and Information Administration (NTIA) describes multiple-choice tests as useful signals, while warning they may not reflect real-world use. A standardized test can make comparisons fairer, but models may respond differently to different interaction styles. Anthropic’s NTIA comment is a company-authored explanation of these trade-offs, not a neutral standard or a current leaderboard.

Can AI benchmark scores be manipulated or misleading?

Scores can be distorted without anyone deliberately falsifying them. The task, test material, setup, or scoring method may produce a result that is real for the test but misleading as a broader claim. The following problems can push a score in different directions:

  • Contamination: If test items or close variants appeared in training data, or became available during testing, a model may benefit from exposure rather than only from the intended capability. For closed models, evaluators may not have access to the full training corpus, which limits direct checks.
  • Broken items: A flawed question, missing file, or ambiguous expected answer can make a capable system appear to fail for reasons unrelated to the intended skill.
  • Reward hacking: A model may exploit a scoring shortcut that earns credit without demonstrating the behavior the test was meant to measure.
  • Refusals and invalid samples: Refusals or compromised test cases can affect results, especially if reports do not explain how such samples were counted.
  • Unreliable graders: An automated scorer can mistake a correct answer for a wrong one, or vice versa. The resulting ranking may reflect grader behavior as well as model performance.
  • Evaluation awareness or underperformance: If a model recognizes an evaluation, or behaves differently for another reason, its score may not represent ordinary use. These possibilities require evidence and should not be assumed from a surprising result alone.

OpenAI’s playbook identifies these as validity concerns, not proof that a particular published score is invalid. A sound report explains which checks were made and how affected items were handled.

What the OpenAI–Anthropic evaluation does—and does not—show

OpenAI described a pilot in which OpenAI and Anthropic ran internal safety and misalignment evaluations on each other’s public models and shared results. In its StrongREJECT v2 section, OpenAI says it selected 60 questions and tested each with roughly 20 variations, including translations and misleading or distracting instructions. OpenAI also cautioned that the range of variations was limited and the automated grader had limitations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Most notably, OpenAI reported that manual review suggested auto-grader errors explained most of an apparent quantitative distinction between some models. That is a concrete reason to check how a score was produced and whether errors were reviewed. It is OpenAI’s account of that exercise, not an independent audit of all claims by either company. The report does not establish a general strategy of deception by OpenAI or Anthropic.

OpenAI’s report on the pilot presents the exercise as a starting point and says additional scaffolding and standardization could make future cross-company evaluations easier. A disclosed limitation is useful context; it does not by itself prove that every other result from the same company is reliable or unreliable.

What contamination studies can tell you

A 2024 NAACL paper, Investigating Data Contamination in Modern Benchmarks for Large Language Models, examines ways to detect possible exposure to benchmark material. It notes that n-gram matching typically depends on having the full training corpus, which is difficult for closed models, while methods that do not require that corpus have their own limits.

One method in the paper masks an unlikely item in benchmark material and asks the model to guess what is missing. Using this missing-option test on MMLU material, the study authors reported exact-match rates of 52% for ChatGPT and 57% for GPT-4. Those figures describe the authors’ particular test and sample. They are not estimates of how much of either model’s MMLU score came from contamination, proof that test answers were intentionally included in training, or evidence that all benchmark results are memorized.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to compare AI benchmark results

Before treating a chart as a ranking, check whether the systems and evaluation conditions are comparable. “Same model name” is not enough if the model version, tools, prompts, or effort allowance differ.

What to check Questions to ask
Task fit What real capability or behavior does the benchmark represent? Is the claim about a tested task, a safety behavior, or performance in a broader workflow?
System and configuration Which exact model version and reasoning setting were tested? What prompts, context, tools, safeguards, interface, and harness were used?
Effort and budget How many turns, tokens, attempts, and retries were allowed? Were wall-clock time, compute, or cost budgets comparable?
Scoring Was success measured by exact match, human review, or a judge model? Were graders checked, partial answers handled fairly, and scoring errors reviewed?
Validity checks Did the evaluator review contamination risk, broken items, reward hacking, refusals, and other reasons a sample might not measure the intended capability?
Uncertainty and replication Does the report state sample size, variation across runs, or uncertainty? Can the result be reproduced under the described conditions?
Practical usefulness Does performance on this test predict the quality, cost, speed, safety, or reliability you need in your own setting?

OpenAI’s playbook for third-party evaluations calls for clear reporting of the system, task, budget, elicitation method, and validity checks. If those details are missing, treat the result as less informative; missing information alone is not evidence of misconduct. For a decision that matters, test the systems on representative examples from your own task and compare the failure modes as well as the headline score.

Why one leaderboard cannot rank every model

A leaderboard compresses a set of choices into a number. Standardized conditions help comparisons, but they may not capture the prompts or interaction styles that work best for a given model, nor the conditions under which a reader will use it. A multiple-choice knowledge result can be informative about those questions without answering how well a system handles a complex, open-ended task.

There is no single definitive method in these sources for ranking all AI systems across every use. The most useful comparison is tied to a defined task, reports enough of its setup to assess fairness, and measures what matters in the intended use. A score is a starting point for that judgment, not a substitute for it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.