iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Two software harnesses can produce very different benchmark scores from the same Claude model. A September 2026 article by Robert Imbeault reports that Backboard CLI scored 85.4% ± 0.8% on Terminal-Bench 2.1 with Claude Opus 4.8, while Claude Code was listed at 78.9%—a 6.5-percentage-point gap. Those are author-reported figures, not an independently verified head-to-head result, and they do not establish that one harness is generally better.
What the Terminal-Bench comparison reports
In his September 18, 2026 DEV Community article, Robert Imbeault attributes the 85.4% ± 0.8% Terminal-Bench 2.1 result to Backboard CLI using Claude Opus 4.8 through Amazon Bedrock. He compares it with a published 78.9% score for Claude Code. The difference is 6.5 percentage points, not a 6.5% relative improvement. The article reports 89 benchmark tasks, five attempts per task, or 445 trials in total.
The figures should be read as a reported comparison, not proof of a controlled experiment. The article is the source for both scores; its underlying leaderboard entries could not be independently verified. The result also describes one benchmark, model version, provider, harness configuration, and evaluation period. It cannot show how the tools will perform on a different model, task set, or everyday workflow.
Imbeault also reports a $280.72 cost for the Backboard CLI run and compares it with $552.67 for a then-verified leader scoring 83.8%. These are date-sensitive, source-reported run costs, not current prices or a like-for-like cost guarantee for other users. Cost comparisons depend on what usage was counted and the evaluation setup.
#1 Best Overall
Why the harness changes the result
A model name alone does not describe the system that completes a benchmark. A harness—the software that runs and coordinates an agent—can shape how the model receives tasks, uses tools, manages context, handles errors, and decides whether to retry. A benchmark score therefore reflects the complete setup, not just the underlying model.
That distinction is supported by a May 11, 2026 Synopticon Research working paper. Its analysis of public results found a median absolute gap of 15.6 percentage points across 64 same-model harness pairs on nine agentic benchmarks. This is an assembled set of benchmark comparisons, not a universal estimate of harness effects in production. Results remain tied to the benchmark, model version, and configurations in each pair.
Rank #2
Other reported results point in different directions
The Synopticon paper reports that Claude Opus 4.5 scored 42.2% on CORE-Bench Hard with Princeton’s CORE-Agent and 77.8% with Claude Code. This is a separate benchmark and an earlier model generation; it is not a replication of the Terminal-Bench 2.1 comparison.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →A GitHub-hosted report describes one Rails-generation task in which Claude Opus 4.7 under opencode had better API correctness and lower reported cost than the Claude Code runs it tested. Its authors caution that the prompt and task were narrow. Taken together, these examples show why a score difference is not a durable ranking: the apparent winner can change with the task and conditions.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to compare harnesses fairly
If you are choosing an agent setup, test both harnesses on the work you actually need done. Keep the underlying model and benchmark tasks aligned, and document the conditions that could change the outcome.
- Match the model and provider. Record the exact model version and provider for each run. A model or provider change makes it harder to attribute a result to the harness.
- Use the same tasks and benchmark version. Keep the task set fixed and identify the benchmark edition. Do not compare scores from different task collections as if they were a direct contest.
- Document the harness configuration. Record prompts, available tools, context strategy, retries, recovery behavior, and other settings that affect the agent loop.
- Run repeated attempts and report spread. Give the number of attempts and the score distribution or uncertainty, not only the best result. Include representative failure cases so readers can see what the aggregate hides.
- Compare costs on equivalent terms. State what the accounting includes and the run period. Report cost alongside score rather than assuming a more expensive run buys a larger improvement; Synopticon found only a weak correlation between cost and score difference across 43 pairs with cost data.
Even a careful benchmark comparison may not predict production performance. Public benchmark results can reflect systems tuned for the benchmark, while real work brings different inputs, tools, and failure costs. Treat the score as evidence about the tested configuration, then validate it on representative tasks before making a deployment decision.
Quick Recap
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

