Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Finding a vulnerability and recognizing that it has been fixed are different skills. In the Attacker-Reachable Sink Triage (ART) benchmark, all seven tested models caught every vulnerable example in the reported run, yet some still mislabeled patched code as vulnerable. That small experiment is a useful diagnostic of patch recognition—not a broad ranking of today’s AI models.
Why vulnerability detection is only half the test
A model can correctly identify an unsafe code path and still overstate risk when a valid security control is present. In a code review, that distinction matters: a false alarm on fixed code can waste review time, obscure real issues, and make findings harder to trust.
ART frames the two questions separately: “did you find a bug?” and “did you respect the fix?” Its author puts the distinction this way: “A model that labels the second snippet reachable_vuln isn’t a worse detector — it’s a worse patch reader.”
How the ART benchmark tests patch recognition
Minimal vulnerable-and-patched twins
The benchmark uses synthetic minimal pairs: two snippets with the same general function shape and identifiers, but a security control is changed between the vulnerable and patched versions. The prompt supplies the code snippet and language, while hiding twin IDs, labels, and rationales. The aim is to focus the test on whether a model notices the change rather than recognizes a memorized vulnerability write-up.
#1 Best Overall
For example, a PHP SQL-injection pair contrasts raw concatenation of attacker-controlled SQL with a version that casts input and uses a prepared statement. The author says the patterns are intended to resemble WordPress-plugin-style PHP and Flask- or Django-request-style Python.
Three tasks, with label triage as the headline
art-label-triage: assign one of four labels:reachable_vuln,patched,safe, orvacuous_noise. Its composite score weights vulnerable accuracy at 0.4, patched accuracy at 0.4, and filler accuracy at 0.2. The author identifies this as the headline metric.art-overconfidence-trap: decide whether patched twins contain a confirmed exploit; the gold answer is no.art-proof-marker-poc: score a minimal lab proof-of-concept marker as 1.0 or 0.0.
The reported dataset contains eight vulnerable/patched pairs spanning SQL injection, cross-site scripting, authentication bypass, command injection, path traversal, local file inclusion, and insecure deserialization across PHP and Python, plus six safe or vacuous controls. That is enough to expose particular response patterns, but not enough to establish how a model will behave across real repositories or vulnerability classes.
What the reported v6 results show
In the author’s art-label-triage v6 run, every one of the seven models found all eight vulnerable twins, yielding 1.000 raw vulnerable accuracy. Their reported scores differed on patched examples and controls:
| Model | ART | Raw vulnerable | Patched | Controls | Twin Gap | Cost (USD) | Latency |
|---|---|---|---|---|---|---|---|
| gemini-2.5-pro | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.181 | 7.9 s |
| gemini-3.5-flash | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.108 | 2.9 s |
| gemini-3.7-flash | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.028 | 9.1 s |
| gemma-4-31b-it | 1.000 | 1.000 | 1.000 | 1.000 | 0.000 | 0.007 | 12.3 s |
| claude-sonnet-4-5-20250929 | 0.950 | 1.000 | 0.875 | 1.000 | 0.125 | 0.060 | 3.1 s |
| claude-haiku-4-5-20251001 | 0.850 | 1.000 | 0.625 | 1.000 | 0.375 | 0.020 | 1.7 s |
| gpt-5.4-nano-2026-03-17 | 0.817 | 1.000 | 0.875 | 0.333 | 0.125 | 0.004 | 1.3 s |
These are the author’s reported results for that task run; the article says the ranked figures come from task-run rewards.score and the table, not the Kaggle collection chart. Model names, prices, and latency are version- and date-sensitive, and the figures should not be read as current general performance or as an independent replication.
Reading Twin Gap and the small denominator
The author defines Twin Gap as vulnerable accuracy minus patched accuracy. A value of zero means equal accuracy on the two kinds of examples; a positive value means the model over-flagged patched examples. Haiku’s reported 0.375 gap corresponds to three misclassified patched examples out of eight. With only eight patched twins, one miss changes patched accuracy and the gap by 12.5 percentage points.
The author reports an exact sign-test p-value of 0.25 for those three misses. That result, alongside the tiny sample, is a reason not to treat the table as a statistically robust large-scale ranking.
What the examples and quality checks add
Some apparent model errors were label errors
The author reports that all seven models disagreed in the same direction with two original labels. Adjudication found the models right: an escaped-input filler was relabeled patched, and a deserialization example that replaced pickle.loads with json.loads was relabeled safe. The initial labels had capped scores at 0.917; after adjudication, the top cluster reached 1.000.
Free tools Windows power users keep installed
One-click scans. No signup required.
This is a concrete reminder that a benchmark score depends on its gold labels as well as the model. If a key is wrong, a model that correctly identifies the code’s behavior can appear to fail.
Interpret individual misses cautiously
The author describes two Haiku misses: a path-traversal twin where the model allegedly ignored basename("../../../etc/passwd"), and an authentication twin where it acknowledged current_user_can but still labeled the example vulnerable because of another risk. These are the author’s interpretations of the examples, not independently tested findings.
Rank #4
The author also reports a Sonnet proof-marker score of 0.0 across retries after a provider returned an empty completion (86 prompt tokens and an empty message). That episode illustrates why a single task cell should be checked against its transcript before it is interpreted as a model capability result.
In additional probes, a red-team persona did not systematically increase overclaiming, and requiring a forced data-flow chain of thought did not eliminate Haiku’s overconfidence-trap error; its reported score moved from 0.625 to 0.50.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsWhat ART can—and cannot—establish
The paired design makes one distinction visible: finding an unsafe example does not prove that a model can consistently recognize a valid fix. But synthetic twins are a narrow diagnostic, not a substitute for testing on diverse, real codebases, where controls may be incomplete, interact with other paths, or depend on surrounding application logic.
Best Value
There is also a design question worth testing in future versions: can a model pass by spotting a familiar fix-like token without reasoning about reachability or whether the fix closes every vulnerable path? A DEV Community commenter suggested adding decoys that contain fix-like tokens while retaining a vulnerable path. That is a proposed extension, not a demonstrated flaw in ART.
The benchmark’s practical contribution is therefore its evaluation framing: report vulnerable detection, patched-code recognition, and controls separately rather than allowing a perfect vulnerability-detection score to stand in for all three. Its reported results support that distinction, while the small synthetic sample limits conclusions about which model is generally best.
Quick Recap
Read the benchmark and source
- The DEV Community article by unit life describes ART and reports the run.
- The article links to the Kaggle Benchmarking Challenge collection, the ART task pages, and the
mziqudhd92/kaggle-art-benchmarksource repository, described there as MIT-licensed. Current program status and availability were not established by the source.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

