Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesIn Erik Hill’s model-drift harness, two short-answer tasks tagged as factual recall could mark a response wrong even when it contained the correct fact. The issue Hill reports is the combination of prompt wording and strict exact-match scoring—not evidence that models lacked the knowledge. He puts the point plainly: “The suspect is the probe, not the models it measures.”
What the harness counted as a failure
Hill describes fact-element with the prompt: “What is the chemical symbol for gold? Two letters only.” The grader lowercased the response and trimmed whitespace, then compared it exactly with au. According to Hill, answers containing the correct symbol could still fail if they added a sentence or markdown emphasis.
For fact-planet, Hill reports that GPT-4o mini returned Mercury. and was marked wrong against the expected short answer. The grader’s exact-match rule treated the added punctuation as a mismatch.
Why the prompt comparison matters
Hill contrasts those tasks with fact-capital, which instructed models: “Answer with only the city name.” In the reported comparison, the models returned Tokyo exactly in 59 and 60 runs, with no failures.
Recommended Free Tools
#1 Best Overall
- Used Book in Good Condition
Hill suggests that “Two letters only” or “One word” might not communicate an output-only constraint as clearly as an explicit instruction. That is his hypothesis, not an established explanation of how models interpret those phrases. What the reported examples show is narrower: some responses included extra text or punctuation, and the grader rejected them.
How much the scoring affected Hill’s suite
Hill reports 60 failures in 60 Claude Sonnet 5 runs, 59 in 59 GPT-4o mini runs, and 23 in 23 Llama 3.1 8B runs for fact-element. He also reports that the task failed on 46 separate days across at least three providers. These are figures from Hill’s 2026 article about his own harness, not an independent benchmark study.
Rank #2
- Author: Karnow, Stanley.
- Publisher: Penguin Books
- Pages: 784
- Publication Date: 1997-06-01
- Edition: 2
The harness has 35 mechanically graded tasks. In Hill’s calculation, the smallest single-task score increment is 100/35, or 2.86 points; fact-element and fact-planet together account for 5.7 points. Hill says the probes also affected the suite’s factual-recall category breakdown. So someone reading that category as a measure of knowledge could, in this particular suite, be comparing response format as well.
The project repository describes the suite’s task structure and capability tags, but that documentation is not an independent verification of the historical run counts reported in the article: project repository.
Rank #3
What the results do—and do not—establish
Hill calls the results solid while describing their cause as uncertain. He reports 26 failures in 49 fact-element runs for Grok 4.5, and says Claude Sonnet 5’s fact-planet behavior differed between historical runs and a later live call. Those observations caution against treating the pattern as a fixed model “house style.”
The article also reports nine live calls through Hill’s probe path. That small set can prompt an audit of outputs and grading rules, but it cannot establish a universal rule about how models understand “only,” nor show that exact-match graders generally mismeasure factual recall. The claims are specific to Hill’s prompts, grader, suite, and reported runs.
How to audit an exact-match short-answer task
Hill’s practical recommendation for maintainers is: “If you maintain an eval suite with exact-match graders on short answers: check what your passing models actually return, not just whether they passed.” A useful audit puts the full response next to the prompt and grading logic:
- Inspect raw outputs. Look for correct answers accompanied by punctuation, explanations, markdown, or other additions.
- Read the prompt literally. Decide whether it clearly asks for only the answer, or merely describes its expected length or shape.
- Check normalization and expected answers. Record exactly which transformations occur before comparison and what string the grader expects.
- State the intended skill. Decide whether the task measures recall, instruction following, strict formatting, or a combination.
- Make the label and score description match. If success requires both the fact and a particular response shape, describe that requirement rather than presenting the score as a clean measure of knowledge.
Choosing what the task should reward
There is a real trade-off, and Hill does not report testing a final fix. A grader that accepts punctuation or additional text may recognize a correct fact, but it could also hide whether a model followed an explicit formatting instruction. Conversely, retaining exact-match scoring makes the required form part of the measured skill. The right choice depends on what the task is intended to test and how its category is presented.
Best Value
Hill considers retagging the affected tasks or describing them more accurately, while warning that leniency can make ignored instructions invisible. His report supports examining the prompt, raw answer, and grader together; it does not establish which alternative grading rule performs best.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

