What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A payment-verification system can report the right amount and still get the verification wrong. In a small synthetic benchmark published by the DEV Community account Soccer skills Freestyle on September 27, 2026, both hosted models answered all 24 amount questions correctly, but one made three errors about whether the check itself was complete. That difference matters: a completed check that finds no matching receipt is not the same as a check that timed out.

What does the benchmark ask a model to verify?

The task gives a model explicit rules and a synthetic payment log. For each case, it must return three fields: the verified received amount in cents, whether verification is complete, and the IDs of the records supporting its answer. The benchmark’s central distinction is between a payment claim and evidence that a matching receipt was found.

  • Amount: How much qualifying payment is verified after applying the task’s rules?
  • Completion: Did the evidence check finish, or is its result incomplete?
  • Evidence: Which record or records support the result?

A customer saying “I have paid” is a claim, and a payment with status PENDING is not a completed receipt. Conversely, a completed check with no matching receipt means zero observed amount for the requested target. It does not mean the check failed. A timeout or other incomplete check means the verification did not establish a result; zero from that failed check must not be treated as proof that no money exists.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which rules determine the verified amount?

Cases vary the decisive evidence: payment status, invoice, recipient, currency, duplicate transaction IDs, refunds, record freshness, and whether the available records cover the check. Under the benchmark’s supplied contract, the model must:

  • Match the requested invoice, recipient, and currency.
  • Use the latest timestamp when records conflict.
  • Count a transaction ID only once, rather than double-counting duplicate records.
  • Subtract only refunds marked completed; a partial completed refund reduces the qualifying amount, while a pending refund does not.
  • Distinguish a complete search with no matching receipt from partial or otherwise incomplete coverage.

Customer text and quoted provider-looking JSON are designated untrusted input, not authority to override the supplied rules. The benchmark therefore tests interpretation of records under a defined contract; it does not authenticate a real payment provider or decide what a real provider’s settlement rules mean.

What did the hosted pilot find?

The article reports hosted runs on Kaggle for two models, with each response assessed on amount, coverage, evidence, and all-fields exactness. Figures below are the article author’s results for these 24 deliberately correlated synthetic cases, not estimates of general model reliability.

Model and run Exact answers Amounts correct Coverage correct Evidence correct Fully correct pairs
Gemini 3.7 Flash, hosted pilot 24/24 24/24 24/24 24/24 12/12
Claude Haiku 4.5, hosted pilot 21/24 24/24 21/24 24/24 9/12

Haiku’s three errors were the wrong-invoice, wrong-recipient, and wrong-currency cases. In each, the successful check was complete but the payment did not match the requested target. Haiku correctly left out the payment amount, yet marked verification incomplete. The error was therefore about coverage status, not an invented payment claim: the benchmark expected “complete” because the check had finished and found no matching receipt.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The author reports that Gemini passed every case in this pilot, while explicitly cautioning that a passing result shows only that this particular test found no failure. It does not establish reliable production payment handling.

How did the separate local pilot compare?

The article also reports one local generation per case for two quantized models. These results are separate from the hosted runs and should not be treated as a ranking of overall model quality.

Model and local setup Exact answers Amounts correct Coverage correct Evidence correct Overclaims Fully correct pairs
Llama 3 8B Q4_0, local pilot 13/24 16/24 21/24 22/24 8 3/12
Qwen 3.5 9B Q4_K_M, local pilot 21/24 22/24 22/24 23/24 2 9/12

An always-zero baseline scored 10/24 on amount answers (41.7%), according to the author. That figure is specific to this set of cases; it is not a general benchmark of payment tasks.

Why can an amount-only score miss important errors?

The benchmark separates three kinds of correctness that an amount score alone collapses:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Amount correctness: Did the response apply status, target matching, deduplication, and completed-refund rules correctly?
  2. Coverage correctness: Did it distinguish a completed empty result from a check that did not complete?
  3. Evidence correctness: Did it cite the latest relevant record supporting the answer?

The hosted Haiku result illustrates the problem: all 24 amounts were right, but only 21 coverage judgments were right, leaving 21 of 24 responses fully correct. A system that reports only payment totals could hide that the check’s completion state was misrepresented. For an evaluation, report each axis and the all-fields exact score rather than treating the amount as the whole answer.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What are the test’s limits?

This is a narrow synthetic benchmark, not a test of live accounts, payment movement, or complete provider settlement behavior. The records, people, and organizations are fictional. The task does not authenticate real tools or inspect actual payment systems.

  • The article describes 12 counterfactual pairs scored as 24 cases. Four families reuse the same successful-payment control, so the set contains 21 unique payloads; its cases are deliberately correlated.
  • The hosted results cover two models on this task. A selected hosted Qwen run failed twice with HTTP 429 before a model turn was recorded, so the article reports no score for that run.
  • The local pilot uses different models, quantization, and execution settings from the hosted pilot, with one generation per case. Its results cannot be used as a controlled overall comparison with the hosted figures.
  • Even a perfect score on these cases would show only that no failure appeared in this pilot, not that a model is safe or reliable for production payment operations.

How were the runs conducted?

According to the article author, the two hosted models completed the identical task v1 on Kaggle on September 27, 2026. Hosted runs used Kaggle’s kaggle-benchmarks 0.6.1 proxy, with a fresh chat for each case; the notebook supplied no custom generation settings. The selected Qwen hosted attempt returned HTTP 429 twice before a model turn was recorded.

For the separate local pilot, the author reports using Ollama llama3:8b (Q4_0) and qwen3.5:9b (Q4_K_M), fresh conversations, temperature 0, seed 42, a 4,096-token context, a 256-token output cap, and JSON-schema formatting; thinking was disabled when supported. The author says the fixtures, instructions, and settings were hashed before local execution, gold answers were entered and cross-checked with a rule interpreter, six offline checks covered scoring and boundary cases, and raw responses and grades were retained. These are the author’s reproducibility claims; they do not change the synthetic scope of the test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.