Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Three AI models all scored perfectly on Jared Chu’s first version of a fictional website-outage benchmark. That result showed the test had hit its ceiling, not that the models were equivalent. In the follow-up, Chu changed one observation at a time in matched scenario pairs. Two models again scored 36 of 36. Claude Haiku 4.5 scored 33 of 36, and its three misses were different kinds of errors, so they should not be counted as one failure rate.

Why the first version told us very little

The first benchmark used five fictional incidents covering DNS, TLS, a deployment rollback, backup recovery, and an incomplete outage report. For each one, the model had to choose an action, cite a supporting evidence statement, and give a short explanation. Chu ran three shuffled answer orders, which produced 15 responses per model. All three models scored 15 of 15.

A perfect score on a test like this means the test could not tell the models apart. It says nothing about whether they would respond the same way under harder conditions. That gap is what the follow-up was built to close.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Changing one fact at a time

The follow-up used six authored pairs. In each pair, the incident description and the list of available choices stayed the same. Only one observation changed, and that change altered both the keyed action and the keyed evidence statement. The question for the model was whether that single fact should change what happens next.

The six pairs test these observations:

  • Whether a prior image passed a compatibility test against the current database schema.
  • Whether DNS tests isolated DNSSEC validation.
  • Whether the errors came from the cache or from the origin.
  • Whether a backup had been validated.
  • Whether a DNS change had been approved.
  • Whether a queued job was durable.

Each variant ran in a fresh conversation. Matched variants used the same option positions across three shuffled orders. That makes 36 responses per model: six pairs, two variants each, three orders. These are 36 responses to six authored pairs, not 36 independent incidents.

The cases and the deterministic scorer were frozen before the follow-up calls. The follow-up was still shaped by what the pilot had revealed, so it is not an untouched holdout test.

How responses were scored

A response earned one point only if it followed the exact JSON schema and selected both keyed choices. The explanation was kept for reading but was not judged automatically. If an infrastructure error occurred, Chu says the run would be invalidated rather than counted as a wrong answer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

All 108 follow-up responses (36 per model) were retained, matched to their frozen prompts, and rescored locally. Chu reports that the local aggregates match Kaggle’s task results.

The three models were chosen before the pilot: one available model from each of three providers. Chu does not claim that any of them is that provider’s strongest offering. All runs used Kaggle platform defaults with no sampling overrides, on September 24, 2026. The pilot, the saved task reruns, and the follow-up are separate result sets and are not pooled.

Results

Model Follow-up responses correct (of 36) Pairs where both variants were correct, per order (of 18) Pairs correct in all three orders (of 6)
Gemini 3.7 Flash 36 18 6
GPT-5.4 mini 36 18 6
Claude Haiku 4.5 33 15 3

These are results from Chu’s benchmark. They are not external or population statistics. The 91.7% figure for Haiku is simply 33 of 36 on this test, and it should not be read as a general accuracy rate.

What Haiku’s three misses were

The three misses fell in three different pairs and had three different causes. Treating them as one failure would hide the differences.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DNS case: validation not isolated

Haiku correctly recognized that the DNS test had not isolated DNSSEC validation. It then chose to inspect the DS and DNSKEY records first, rather than trace resolution for the key. The evidence reading was right, but the action differed from the keyed one.

Cache case: bypass evidence recognized, origin checked first

Haiku recognized that the cache-bypass evidence was successful. It then chose to check origin health before evicting the cache. Chu says that extra diagnostic step may be defensible, and that this result does not establish unsafe behavior.

Durable-queue case: schema violation

Haiku selected both keyed IDs correctly but added an unrequested reason2 field to its JSON. Because the scorer requires the exact schema, this counted as a miss even though the decisions were right.

In short, two of the misses are action-key disagreements and one is a schema violation. The model’s decisions in the cache and DNS cases were arguably reasonable, while the queue case was a formatting failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the scores do not show

  • The cases were short multiple-choice exercises with explicit runbooks and some easy distractors.
  • The models did not investigate a live outage, execute a change, handle new evidence over time, or demonstrate recovery.
  • The evidence choices measured whether a model recognized appropriately scoped claims. They did not measure general confidence calibration.
  • Six authored pairs cannot establish a general ranking of models.
  • The shuffled orders reveal variability, but identical prompts were not repeated enough to separate option-position effects from sampling variability.
  • No independent expert validation or human manual review was carried out. AI tools drafted the cases, implemented and ran the evaluation, analyzed the outputs, and wrote the article. That makes the answer key a benchmark convention, not verified operational ground truth.
  • All cases were fictional. No customer data or real infrastructure was involved.

Chu suggests that a future version could bring in independent operator review of disputed actions, repeat identical prompts, and stage an incident in which the model must request missing evidence before proposing a change.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Inspecting and reproducing the benchmark

  • The Kaggle project contains separate tasks for the pilot and for the paired follow-up.
  • Kaggle’s model headers may show 0.00. That reflects a “No overall score” display setting, not an additional measured result.
  • The paired notebook publishes the full corpus, answer key, scorer, and run exports.
  • To compare the three models’ traces, open the registered task’s Compare Outputs view.
  • Kaggle Benchmarks SDK handles task registration and model execution. The case content and scoring logic were written for this submission.
  • The work is public under the Apache 2.0 license. Reproduction uses order seeds 11, 29, and 47. The frozen paired corpus and scorer have the SHA-256 hash 0def44fe0c0e9d483487ecaaa0b8a8ccba4a30c8127b02e11e3a91d1eab34295.

Platform views may change, so check the project page against these descriptions before relying on them.

How to read AI incident benchmarks

  • Look past the headline total. Separate the joint score from the stricter measure that requires both variants of a pair to be correct.
  • Classify each miss as an action-key disagreement or a format violation. The two call for different fixes.
  • Check whether the test ever changes a single fact. A test that holds the scenario fixed while varying one observation shows whether a model responds to the evidence.
  • Keep benchmark results separate from claims about safety, live operations, or general model quality.

Chu’s central point is that perfect scores are a warning sign when a test is too easy to separate models, and that the next useful step is a test where one fact can change the right answer.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.