Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Action scaling can outperform rerunning whole agent trajectories when a capable verifier can choose a better next command from several candidates before any of them changes the environment. In one TerminalBench-Lite comparison, Mid-Harness combined with Best-of-3 trajectories reached 66.33% Pass@1, versus 59.18% for Best-of-7 alone, at lower estimated reference-priced token cost. That is evidence for a promising approach in a tested setting—not a rule that action sampling always wins.

How candidate verification works

Mid-Harness adds inference at the boundary between an action-generating model and the execution harness. At each step, the generator proposes several possible next actions from the same interaction history. A verifier compares those candidates, and the harness executes the selected action. The central comparisons keep the generator and harness fixed, changing where additional computation is applied.

This differs from trajectory scaling. Instead of comparing or refining completed task runs, action scaling evaluates alternatives before the chosen command reaches the environment. That timing matters in terminal work: commands can change files, install packages, or otherwise alter the state that later decisions depend on.

For example, the accompanying DEV Community article illustrates how a package-install typo—pip install yaml instead of pip install pyyaml—can send a task down the wrong path. This is an illustration, not a measured result from the paper.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the reported results show

More candidate actions help when verification is strong

In the paper’s central TerminalBench-Lite comparison, the TMAX-9B base agent achieved 50.00% Pass@1. Using the same generator with eight sampled actions and a GPT-5.6 Sol verifier raised the reported score to 68.03%. These are results from that experiment, not a forecast for another model or task.

The authors also report a verifier-distillation comparison in which Pass@1 rose from 54.76% to 57.14% while the action generator stayed unchanged. Their findings indicate that the verifier’s ability to identify a suitable candidate matters: wider sampling offered little benefit with weak verification. Among the evaluated self-verification methods, pairwise verification performed best.

One comparison favors combining the two kinds of scaling

On TerminalBench-Lite with TMAX-9B, Mid-Harness combined with Best-of-3 trajectories scored 66.33% Pass@1, compared with 59.18% for Best-of-7 alone. The combined setting also had lower estimated reference-priced token cost, according to the authors. Those cost figures are experiment-specific estimates—not measured deployment bills or a universal cost-per-task result.

Results vary by benchmark

The paper reports other evaluations, but gains are not identical across tasks or verifier variants. For example, its table reports TMAX-9B rising from 21.72% to 27.34% Pass@1 on Terminal-Bench 2.1 under zero-shot verification, and from 46.67% to 48.00% on the SWE-bench-Verified Mini subset. These examples should not be read as a guarantee of improvement in every setting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why “beats” needs a qualification

The 66.33% versus 59.18% comparison supports the headline for that evaluated configuration. It does not show that action-level scaling is always more successful or cheaper than rerunning trajectories. The authors describe action scaling as complementary to trajectory scaling; in practice, a system might use both.

Verification is also a substantive requirement, not a free consequence of generating alternatives. The authors report no gold action labels, which limits direct measurement of candidate coverage and verification correctness. Their analysis finds persistent disagreement with the stronger verifier about command semantics and execution feasibility, and distillation does not close the full gap to frontier verification. The results therefore do not prove that an arbitrary harness verifier will select commands safely or correctly in production.

How to compare the methods in your own agent

Use the same evaluation conditions when deciding whether to spend inference on action candidates, complete trajectories, or both:

  • Task success: compare the same benchmark, task set, and metric. Do not conflate Pass@1 with Pass@3.
  • Inference cost: identify whether you are comparing token counts, estimated reference-priced cost, or actual deployment spend. The paper’s cost figures are estimates tied to its reference pricing.
  • Verifier method and strength: record whether candidate selection uses a stronger external verifier, self-verification, pairwise comparison, or a distilled verifier.
  • Environment executions: account for the number of complete trajectories actually run. Action filtering can produce a returned run using one environment instance, while trajectory sampling may execute multiple complete runs.
  • Transfer to your task: name the model, benchmark, and harness. The study reports varied outcomes across them, so one benchmark is not a substitute for evaluation in your target system.

For the paper’s method, candidate actions are compared before execution. That design can avoid committing to a weak command when a better alternative is available, but the benefit depends on the verifier recognizing that alternative. A useful local test should therefore measure success, verifier quality, execution count, and cost together rather than treating candidate count alone as the control to optimize.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Source

The primary source is Kang and colleagues’ “Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents”, posted September 30, 2026. Its abstract states: “With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator.”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.