iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Neither the model nor its scaffolding is the universal bottleneck. An agent’s long-running performance depends on the model, the software harness around it, and the task environment together. To find what is limiting a particular system, compare controlled configurations and inspect whether they finish the work correctly—not just whether an automated grader says they passed.
What counts as the model, and what counts as scaffolding?
The model generates reasoning and selects actions. The scaffold, also called a harness, is the surrounding software and evaluation setup: it determines what context the model receives, which tools it can use, how it manages steps or state, what feedback it gets, and how completion is judged.
That division matters because a harness can affect how effectively a model uses its capabilities. Better context, tool access, recovery mechanisms, or verification could improve results; unsuitable versions could hold a capable model back. But scaffolding cannot guarantee sound reasoning or correct execution. These are explanations to test in a given system, not a universal finding that one component always dominates.
Why long-running work needs its own evaluation
A short task may not reveal problems that emerge over many steps: an agent can lose track of earlier requirements, fail to recover from a bad action, or stop before the work is complete. METR proposes measuring agents by the length of tasks they can complete, while SWE-Bench Pro describes software tasks that may take hours to days. A score on a short benchmark therefore does not, by itself, establish reliable performance on longer work.
#1 Best Overall
METR’s 2025 paper reports that, for its selected tasks and measurement method, the task length frontier AI systems could complete with 50% reliability doubled approximately every seven months over the studied 2019–2024 period. This is a historical estimate for that dataset and method—not a forecast, a guarantee of future progress, or evidence that scaffolding alone caused the trend. METR’s long-task measurement and its HCAST task resource cover varied work, including software engineering, machine-learning engineering, cybersecurity, and general reasoning. HCAST tasks have estimated human completion times from one minute to more than eight hours.
What benchmark results can—and cannot—tell you
Benchmark scores belong to tested configurations, not to model names in isolation. OpenAI’s MLE-bench describes its best-performing setup as o1-preview paired with AIDE scaffolding. That setup reached at least the level of a Kaggle bronze medal in 16.9% of competitions in the benchmark. The figure is specific to that pairing and task set; it is not a general success rate for agents.
Rank #2
OpenAI’s SWE-bench Verified discussion likewise treats scaffolding and external enhancements as relevant when assessing capability. Its article gives historical leaderboard context dated August 5, 2024: top agents scored 20% on SWE-bench and 43% on SWE-bench Lite at that time. Those are dated figures, not current scores, and they should not be compared directly with results from a different benchmark suite. OpenAI’s SWE-bench Verified article says: “Community-led progress in agent scaffolding highlights the need to consider potential external enhancements to a model when assessing risk.”
Recommended Free Tools
Automated success checks also need scrutiny. OpenAI’s o1 system card reports that manual inspection of some trajectories that passed an autograder found major portions of the task silently incomplete. A pass condition can therefore overstate completion if it does not check the actual requirements.
Rank #3
How to test whether the model or harness is limiting performance
- Define task-specific success. Specify what counts as correct and complete before comparing systems. Use checks that cover the whole task, not just a convenient proxy.
- Hold the evaluation constant. Keep the task set, environment, tool access, budget, and success criteria the same. Where practical, change one factor at a time—for example, compare models under the same harness, then harnesses with the same model.
- Record the complete configuration. Report the model and scaffold together, along with relevant tools, budget, and environment. A model name alone does not identify the tested system.
- Inspect trajectories. Review whether the agent omitted requirements, used tools appropriately, recovered from errors, and stopped for a sound reason. This can reveal false passes that a grader misses.
- Use varied tasks and report more than pass rate. Include tasks with different lengths and demands. Compare completion, task duration or length, process quality, efficiency, and failure behavior; include time or compute cost when it is measured. Repeat runs when assessing reproducibility.
These dimensions can move in different directions, so a single leaderboard ranking may not identify the best configuration for a particular workload. The cited work supports measuring long-task ability and multiple outcomes, but does not establish one common protocol for every agent domain.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Where the evidence is strongest
The available examples focus chiefly on coding and machine-learning engineering agents. Results depend on the tasks, model, scaffold, budget, environment, and grader used. They support treating performance as a property of the full configuration and evaluating longer tasks directly; they do not establish that the model or scaffold is the primary bottleneck across all agents.
Quick Recap
Best Value
Rank #4
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

