The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →No available evidence shows that Span-01 and Mercury Decide earned the same score or failed in opposite ways. Respan published Span-01 results on behavior-classification benchmarks; a Reddit author reported Mercury Decide results on a narrow Korean-language task about Roblox Terms of Service. Because the tests used different cases and goals—and Span-01 was not tested on the Reddit cases—the results are not a head-to-head comparison.
What Span-01 and Mercury Decide are designed to do
Span-01: classify behaviors in conversational traces
Respan describes Span-01 as a classifier that applies natural-language behavior definitions to conversational traces. It returns a probability for each behavior being present, absent, or not_observable in one forward pass. A developer can combine those probabilities with thresholds and code to alert, block, log, route a case to a person, or send uncertain cases for review. Respan’s launch post describes the model as reasoning over behavior definitions and returning probabilities across them in parallel; its documentation describes the product and its use.
Mercury Decide: answer structured decision questions
Mercury Decide is described as a decision model for Choice, Score, and yes/no questions, with probabilities attached to its answers. The profile describes access through OpenRouter’s System One endpoint and identifies the service as early access. Claims in that profile about a JevBench ranking and throughput of up to 14 decisions per second are attributed to Inception, not independently verified there. The Mercury Decide profile reflects details as of October 1, 2026.
What the published numbers actually measure
The reported figures come from separate datasets, tasks, and evaluation processes. They should be read in their original context—not compared as if they were scores from one contest.
Recommended Free Tools
#1 Best Overall
| System and figure | What it measures | Source and qualification |
|---|---|---|
| Span-01: 0.843 overall F1 | Behavior benchmark; Respan says this overall figure is the unweighted mean of English and multilingual F1. | Respan, 2026. Launch post and documentation. |
| Span-01: 0.806 overall F1 | Respan’s production-behavior benchmark. The same table reports 0.716 for Jev, 0.719 for Sonnet 5, and 0.885 for GPT-6 Sol. | Respan, 2026. This is a separate benchmark from the Korean Roblox report task. Launch post and documentation. |
| Mercury Decide: 66.7% accuracy; 28 false negatives among 90 cases | A Korean-focused test of whether chat logs violate Roblox Terms of Service. | Reddit benchmark author, October 1, 2026. Author-reported result from one task, not a general model ranking. Reddit post. |
Respan also reports a separate evaluation of 11 decision models across accuracy, consistency, injection resistance, and calibration. In that vendor-published evaluation, Jev 1.13.0 scored 0.932 accuracy, a 0.021 paired flip rate, a 0.063 injection attack success rate, and 0.045 expected calibration error. This is not a Mercury Decide result or a comparison with Mercury Decide. Respan’s launch post describes the evaluation; ModelSystem.One notes that Respan’s benchmark labels are model-generated, mostly through agreement between GPT-5.6 Sol and Claude Opus 5, rather than ground truth.
What the Mercury Decide failure report does—and does not—show
The Reddit author reports 28 false negatives in 90 cases and says Mercury Decide appeared to answer “no” on almost every possible report case at the tested threshold. The post limits its finding to understanding Korean and deciding whether chat logs violate Roblox’s rules. It does not report Span-01 results on those same cases. Read the author’s benchmark report.
Rank #2
Respan’s Span-01 material evaluates behavior detection across English and multilingual data, as well as separate production-behavior domains. Its published categories include jailbreak and prompt injection, safety and refusals, privacy and secrets, hallucination and grounding, agent and tool reliability, task and instruction following, and response quality. Those evaluations do not establish how Span-01 would perform on Korean Roblox reports.
That distinction matters because a behavior classifier, a fixed-choice decision model, and a report/no-report workflow do not necessarily have the same task definition or output. Nor does a shared label such as “accuracy” or “F1” make results comparable when the data, labels, class balance, and decision threshold differ.
Free tools Windows power users keep installed
One-click scans. No signup required.
How to make a fair Span-01 vs. Mercury Decide comparison
A credible head-to-head would run both systems on the same cases, with the task and scoring rules set before evaluating either one. It should disclose:
- Task and output: whether each model is classifying behavior in a trace or answering a fixed-choice, score, or yes/no question.
- Cases and labels: the shared dataset, how labels were established, and the balance of positive and negative examples.
- Error counts at a common threshold: report false positives and false negatives alongside accuracy or F1. This is essential for a reporting workflow where missed violations may matter more than unwanted reports.
- Consistency and adversarial behavior: test whether equivalent inputs produce changed decisions and whether injected instructions alter results. Respan includes flip rate and injection resistance in its separate decision-model evaluation; that does not supply Mercury Decide’s corresponding results.
- Calibration: compare predicted probabilities with observed outcomes using a named metric and the same labeled cases. Respan reports expected calibration error in its separate evaluation.
- Language, version, and access route: state the language and use case for each result, the model version, the endpoint used, and the test date. A Korean-specific result should not be presented as general performance.
- Operational terms: compare current endpoint limits, latency, pricing, and hosting terms. Respan lists Span-01 input pricing with free output; the Mercury Decide profile describes a free early-access route but says some limits and paid pricing are unpublished. These details may change, so check the providers’ current materials before making an implementation decision.
How to interpret the comparison today
The defensible conclusion is limited: Respan reports Span-01 performance on its behavior benchmarks, while a Reddit author reports a Mercury Decide result on one Korean Roblox reporting task. The available figures do not establish equal scores, opposite failure patterns, or which system is better overall. Respan’s scores also need the qualification that its benchmark labels are model-generated, as noted by ModelSystem.One.
Quick Recap
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

