Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

On Day 3, Sean Campbell’s benchmark found a sharp weak spot in one model—and Campbell found a similar problem in his own AI-assisted writing workflow: ambiguous evidence had been turned into a confident claim. The lesson is to inspect failures by task, check whether results repeat, and avoid treating an unclear note as a fact.

What Day 3’s benchmark measured

In a 2026 post on DEV Community, Sean Campbell describes a benchmark of 200 invented items divided into four task shapes: route, classify, judge, and ground. One in five items, by his account, could only be answered correctly with ESCALATE. The benchmark tracked task score separately from false confidence: cases where a model answered despite lacking enough evidence.

For Day 3, Campbell compared the weakest task shape—or “floor”—for 12 hosted models, using Wilson intervals. These are the author’s reported results, not independently validated measurements. The post’s central point is that an overall score can hide a serious failure on one kind of task.

Which task shape was weakest?

Ground was the weakest shape for Gemini 3.7 Flash, Gemini 3.1 Pro, Claude Sonnet 5, Claude Opus 5, Gemini 3.8 Flash, GPT-5.5, and GPT-5.4 nano. Classify was weakest for Qwen3 235B Instruct, Claude Haiku 4.5, Gemma 4 26B, gpt-oss-20b, and DeepSeek-R1.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Haiku’s comparison covered only three shapes: all of its route calls failed. That missing shape matters when interpreting its floor, because the result does not represent a complete four-shape comparison.

Why false confidence matters more than a weak average

The most striking result in Campbell’s post is Claude Haiku 4.5’s judge performance: it answered 9 of the 10 judge items that should have been escalated. Campbell reports this as a 90% false-confidence rate on that shape. Across Haiku’s three measured shapes, it answered 10 of 28 unanswerable items anyway; the failures were concentrated in judge rather than evenly distributed across tasks.

That distinction changes what a model score means. A model can do reasonably on answerable cases while still being unsafe for a task that requires recognizing uncertainty. For a model comparison, look at the false-confidence rate on unanswerable items as well as task score, and ask how many such items were tested.

What the floor does not establish yet

A worst-shape view helps expose a weak spot, but small samples make close comparisons uncertain. For the top six rows in Campbell’s table, each shape had only 8 to 12 unanswerable items. Even where no false-confidence failures were observed, the post estimates upper bounds of roughly 24% to 32%; the intervals overlap. A zero in a small sample is therefore not proof that the underlying risk is zero, and the reported results do not support a confident ranking among those models.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Campbell also compared repeat runs of four frontier models across the same 200 items. The figures below are same-answer counts from two runs, as reported in the 2026 post:

Model Same answer across two runs
Claude Opus 5 199/200 (99.5%)
Claude Sonnet 5 195/200 (97.5%)
Gemini 3.1 Pro 195/200 (97.5%)
GPT-5.5 194/200 (97.0%)

The intervals overlap, so these counts do not establish which model is more consistent. They describe agreement on this benchmark under the reported run conditions, not a general reliability ranking.

When an apparent model failure may be an evaluation artifact

Campbell says Gemini’s five verdict changes came from replies that hit an output-length cap and parsed successfully in only one run—not from different substantive answers. The scorer treated a parsing error as its own verdict. That means a score can reflect a pipeline’s handling of capped or malformed output as well as the model’s answer.

The runs also did not use identical generation settings: only Gemini ran at temperature 0; the two Claude 5 models rejected that setting, and GPT-5.5 used its default. In addition, Campbell said the third run for classify, judge, and ground had hit Kaggle’s daily spend cap and was expected the following day, so the reported figures were not yet final.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Check output handling: determine whether a response was truncated, and whether a parser error is scored as a model verdict.
  • Check settings: record temperature and other generation parameters, and note when models cannot use the same configuration.
  • Check run completeness: confirm that the expected items were collected and distinguish preliminary from completed comparisons.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What Campbell learned about Kaggle runs and retries

Campbell cautions that Kaggle’s run timer did not appear to match the time spent making calls: he observed repeat runs taking 2–5 seconds for 40–60 items, while downloads contained all expected items. Those timings are his observations, not a verified description of Kaggle’s platform behavior.

He also describes a five-minute sandbox task that retried, was killed at 300 seconds, and resubmitted paid runs, creating duplicate spend. His proposed safeguard is to separate submission from collection and make paid actions refuse duplicate runs:

  1. Submit paid runs in one short task.
  2. Collect the completed results in a separate task.
  3. Make submission logic detect an existing run and refuse to launch a duplicate paid action.

The benchmark caught its author, too

The title refers to Campbell discovering that an earlier AI-assisted writing session had misread a terse note as a grade. He had not graded anything, but the session recorded a grade in his voice, and he published it without noticing. His correction is straightforward: if a note might be a grade, preserve its wording and ask what it means instead of silently resolving the ambiguity.

That personal mistake mirrors the benchmark’s central concern. A system should not turn unclear evidence into a more certain claim than the evidence supports. As Campbell puts it: “That’s the benchmark’s whole subject, happening one level up: an answer stated with more confidence than the evidence behind it, by a system that should have said "I’m not sure what you meant."”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to use these results in a model comparison

Campbell’s Day 3 report is most useful as a reminder to evaluate more than a headline score. When comparing models on a benchmark like this, examine:

  • the weakest task shape, not just the aggregate result;
  • false-confidence rates on items that require escalation, alongside the number of those items and interval width;
  • agreement across repeated runs, without treating overlapping intervals as a ranking;
  • how output caps, parsing errors, timers, incomplete runs, and retries affect the recorded result; and
  • whether generation settings are comparable across models.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.