The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →On Day 3, Sean Campbell’s benchmark found a sharp weak spot in one model—and Campbell found a similar problem in his own AI-assisted writing workflow: ambiguous evidence had been turned into a confident claim. The lesson is to inspect failures by task, check whether results repeat, and avoid treating an unclear note as a fact.
What Day 3’s benchmark measured
In a 2026 post on DEV Community, Sean Campbell describes a benchmark of 200 invented items divided into four task shapes: route, classify, judge, and ground. One in five items, by his account, could only be answered correctly with ESCALATE. The benchmark tracked task score separately from false confidence: cases where a model answered despite lacking enough evidence.
For Day 3, Campbell compared the weakest task shape—or “floor”—for 12 hosted models, using Wilson intervals. These are the author’s reported results, not independently validated measurements. The post’s central point is that an overall score can hide a serious failure on one kind of task.
Which task shape was weakest?
Ground was the weakest shape for Gemini 3.7 Flash, Gemini 3.1 Pro, Claude Sonnet 5, Claude Opus 5, Gemini 3.8 Flash, GPT-5.5, and GPT-5.4 nano. Classify was weakest for Qwen3 235B Instruct, Claude Haiku 4.5, Gemma 4 26B, gpt-oss-20b, and DeepSeek-R1.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
Haiku’s comparison covered only three shapes: all of its route calls failed. That missing shape matters when interpreting its floor, because the result does not represent a complete four-shape comparison.
Why false confidence matters more than a weak average
The most striking result in Campbell’s post is Claude Haiku 4.5’s judge performance: it answered 9 of the 10 judge items that should have been escalated. Campbell reports this as a 90% false-confidence rate on that shape. Across Haiku’s three measured shapes, it answered 10 of 28 unanswerable items anyway; the failures were concentrated in judge rather than evenly distributed across tasks.
That distinction changes what a model score means. A model can do reasonably on answerable cases while still being unsafe for a task that requires recognizing uncertainty. For a model comparison, look at the false-confidence rate on unanswerable items as well as task score, and ask how many such items were tested.
What the floor does not establish yet
A worst-shape view helps expose a weak spot, but small samples make close comparisons uncertain. For the top six rows in Campbell’s table, each shape had only 8 to 12 unanswerable items. Even where no false-confidence failures were observed, the post estimates upper bounds of roughly 24% to 32%; the intervals overlap. A zero in a small sample is therefore not proof that the underlying risk is zero, and the reported results do not support a confident ranking among those models.
Campbell also compared repeat runs of four frontier models across the same 200 items. The figures below are same-answer counts from two runs, as reported in the 2026 post:
| Model | Same answer across two runs |
|---|---|
| Claude Opus 5 | 199/200 (99.5%) |
| Claude Sonnet 5 | 195/200 (97.5%) |
| Gemini 3.1 Pro | 195/200 (97.5%) |
| GPT-5.5 | 194/200 (97.0%) |
The intervals overlap, so these counts do not establish which model is more consistent. They describe agreement on this benchmark under the reported run conditions, not a general reliability ranking.
Rank #4
When an apparent model failure may be an evaluation artifact
Campbell says Gemini’s five verdict changes came from replies that hit an output-length cap and parsed successfully in only one run—not from different substantive answers. The scorer treated a parsing error as its own verdict. That means a score can reflect a pipeline’s handling of capped or malformed output as well as the model’s answer.
The runs also did not use identical generation settings: only Gemini ran at temperature 0; the two Claude 5 models rejected that setting, and GPT-5.5 used its default. In addition, Campbell said the third run for classify, judge, and ground had hit Kaggle’s daily spend cap and was expected the following day, so the reported figures were not yet final.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
- Check output handling: determine whether a response was truncated, and whether a parser error is scored as a model verdict.
- Check settings: record temperature and other generation parameters, and note when models cannot use the same configuration.
- Check run completeness: confirm that the expected items were collected and distinguish preliminary from completed comparisons.
What Campbell learned about Kaggle runs and retries
Campbell cautions that Kaggle’s run timer did not appear to match the time spent making calls: he observed repeat runs taking 2–5 seconds for 40–60 items, while downloads contained all expected items. Those timings are his observations, not a verified description of Kaggle’s platform behavior.
He also describes a five-minute sandbox task that retried, was killed at 300 seconds, and resubmitted paid runs, creating duplicate spend. His proposed safeguard is to separate submission from collection and make paid actions refuse duplicate runs:
- Submit paid runs in one short task.
- Collect the completed results in a separate task.
- Make submission logic detect an existing run and refuse to launch a duplicate paid action.
The benchmark caught its author, too
The title refers to Campbell discovering that an earlier AI-assisted writing session had misread a terse note as a grade. He had not graded anything, but the session recorded a grade in his voice, and he published it without noticing. His correction is straightforward: if a note might be a grade, preserve its wording and ask what it means instead of silently resolving the ambiguity.
That personal mistake mirrors the benchmark’s central concern. A system should not turn unclear evidence into a more certain claim than the evidence supports. As Campbell puts it: “That’s the benchmark’s whole subject, happening one level up: an answer stated with more confidence than the evidence behind it, by a system that should have said "I’m not sure what you meant."”
How to use these results in a model comparison
Campbell’s Day 3 report is most useful as a reminder to evaluate more than a headline score. When comparing models on a benchmark like this, examine:
Quick Recap
- the weakest task shape, not just the aggregate result;
- false-confidence rates on items that require escalation, alongside the number of those items and interval width;
- agreement across repeated runs, without treating overlapping intervals as a ranking;
- how output caps, parsing errors, timers, incomplete runs, and retries affect the recorded result; and
- whether generation settings are comparable across models.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

