Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Before you publish a coding-agent leaderboard, freeze three things: the identity of the task pack, the scoring rules and their version, and the negative controls that show the grader can reject bad work. Then publish the ranking together with those three items. A bare score tells a reader what one run produced. It does not tell them whether candidates were compared on the same tasks under the same rules.

The method described here comes from a DEV Community article by Avery Wang. The search listing showed a September publication date but no year, so none is given here. Wang’s core position is that “a published agent score is trustworthy only after the dataset, metric functions, and control runs are frozen.” The sections below set out what that means in practice, and where the proposal stops short of proof.

What to freeze before ranking candidates

The leaderboard is only interpretable if three inputs are fixed and identifiable before any candidate is scored: the task pack, the metric, and the controls. Each one can change silently in ways that alter the comparison, so each needs an explicit identity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

1. The task pack

Hash every file in the benchmark and write a manifest that lists the hashes. Any later change to a task, fixture or test then produces a different pack identity. The article’s programming-task setting also calls for two things:

  • Hidden tests that candidates cannot read or edit during the run.
  • A declared per-item time budget, written down before scoring starts.

If a pack turns out to be flawed, the article’s advice is to publish a new pack identifier rather than quietly altering the old one. The old ranking then stays tied to the pack it was produced from, and the corrected pack is scored as a separate comparison.

2. The metric

Define each outcome as an artifact or an observable event. Report the component outcomes separately: compile, test pass, lint, timeout, and assertion deletion. Keeping these visible matters because an opaque blended score can hide a candidate that passes tests by deleting the assertions that would have failed it.

If you still want one headline number, the article’s example uses a strict conjunction: a task counts as solved only when every required component succeeds. Whatever formula you choose, record its version alongside the pack identifier.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. The control battery

Negative controls are inputs that should score as failures. The article’s sample battery includes three:

  • An empty patch, which changes nothing in the repository.
  • Shuffled tests, where test files are paired with the wrong tasks.
  • An echoed prompt, where the model’s output is just the task text.

If a control passes, the grader or the pack is at fault, not the candidates. Controls are what show that the harness can say no.

4. The report that travels with the ranking

The article argues that a number without its manifest, metric and controls is not an adequately documented measurement. A complete report therefore includes the pack identifier, the metric version, the control results, and each candidate’s component outcomes, not only a headline rank.

How to run the control gate

The article proposes a publish gate that refuses a leaderboard when controls pass unexpectedly. In practice the sequence is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Run each negative control through the same harness, with the same time budget and sandbox, that candidates will use.
  2. Check that every control scores as a failure on every component outcome.
  3. If any control passes, stop. Do not publish a ranking. Find out whether the task pack, the hidden tests or the grader is wrong.
  4. If you fix the cause, treat the corrected inputs as a new pack identifier and rerun the controls before any candidate is scored.

The gate only blocks publication. It does not tell you which part of the setup failed, so the investigation in step 3 is still your job.

What freezing does not guarantee

A frozen manifest preserves identity. It proves that the same tasks and rules were used. It does not prove that those tasks or rules were good. The article itself lists limits that a frozen pack cannot fix:

  • Task relevance. Frozen tasks may not represent the work you care about.
  • Test quality. Hidden unit tests are a poor oracle for interface work, migrations and incident response when fixtures fail to encode the real loss.
  • Quality dimensions. The method does not measure taste, architecture or long-horizon refactors.
  • Single-number claims. The article warns against using the method as a single promotional percentage.
  • Execution safety. Teams need an execution sandbox before running model-authored patches.

A fixed pack and a passing control gate are necessary conditions for a fair comparison. They are not sufficient ones.

Private pack or public benchmark

There are two broad routes. You can build and freeze a private task pack that matches your own work, or you can rely on an existing public, versioned benchmark that already ships hidden tests and documented controls. The article does not compare named commercial benchmarks, and it provides no head-to-head numbers, so the table below lists the criteria to check rather than verdicts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Criterion Private frozen pack Public versioned benchmark
Task relevance Written for your claim and codebase; you must justify each task Set by the maintainers; check whether it matches the work you are assessing
Version identity Your own manifest and hashes Whatever versioning the maintainers publish; confirm the exact version you cite
Grader transparency Fully visible to your team, which also means you must audit it yourself Depends on what is published; check whether the grader and hidden-test policy are documented
Leakage risk Low if tasks are not published; review who has seen them Check whether tasks or tests are public, since public material can end up in training or prompts
Control coverage Only what you build; the article’s three controls are a starting point Check whether documented controls exist and whether they are the ones you need
Coverage of quality dimensions in your claim Limited to the tasks you wrote Limited to what the benchmark measures; compare against your claim

A checklist for reading someone else’s agent score

When you are judging a published ranking, including your own, look for these items before trusting the order:

  • A pack identifier, with a statement of whether it is the same pack used for every candidate.
  • A metric version, and the component outcomes behind any headline number.
  • Negative control results showing that controls scored as failures.
  • Time budgets and the execution environment, stated in the report.
  • A match between the tasks and the claim being made. A score on bug-fixing tasks does not establish performance on architecture work.

If any of these is missing, the ranking may still be informative, but it cannot be checked or repeated.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Freezing outside coding agents

Other benchmark projects make the same choices explicit, although they do not validate the article’s method. A public ARC-AGI-3 project plan describes hashing and archiving a fixed evaluation manifest, using the game as the unit of generalization, aggregating repeated seeds within a game, and predeclaring how crashes, timeouts and missing results are handled. An earlier version of that plan stated that the fixed evaluation set and aggregation matter to the result, including treating untouched games in the relevant leaderboard split as zero. That is a domain-specific choice, not a general rule.

An autonomous-vehicle policy lab decision log shows freezing being made conditional. The team delayed freezing after validity concerns, and later froze scenario sets with hashes and a disjoint selection probe. The lesson transfers in one direction only: a freeze is worth doing after the design problems that would make it meaningless have been resolved.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the evidence does and does not establish

The article’s scripts, gate table and workflow sketches are proposed and unexecuted. The source reports no trial of its own harness and no measured benchmark results. Its numeric values, including sample thresholds and rates, are illustrative criteria, not findings to quote. The article also discloses that it was prepared as part of MonkeyCode product outreach, so its recommendations are best checked against independent sources before you adopt them.

What the article establishes is a clear argument: a leaderboard score depends on the identity of the tasks and the scoring rules, so both should be fixed, versioned and published with the result. That argument stands on its own logic. Whether this particular control battery catches every failure mode is still open.

Finally, the article’s examples do not establish a universal standard for evaluating agents across all kinds of work. Treat them as a protocol to adapt and test on your own tasks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.