What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A strict stability filter can lose statistical power as you add repetitions, even when the models under comparison have not changed. In Erik Hill’s account of a 159-task evaluation suite, a rule that discarded any task whose configuration disagreed with itself produced a significant paired result on three repetitions, then a weaker result once the discard count grew and the informative count shrank after more data were pooled. The lesson is about the filter and how its retained sample changes, not about which model is better.

What the strict rule does

Hill compared two models on a frozen suite of 159 tasks. Each task was run across repeated trials, and the suite is built from trap questions. His first rule, strict, discarded a task if a single configuration disagreed with itself across repetitions. Only the tasks that survived were treated as informative and fed into a paired comparison between the two models.

The filter is attractive for a simple reason: a task whose answer flips between runs is noise, and counting it as a win or loss inflates confidence. The trade-off is that the filter also removes data, and the amount it removes depends on how many repetitions you run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The first result, and the rule written after it

Across repetitions 1 to 3, strict produced a 7–1 comparison over 8 informative tasks, with p=0.070. The filter had discarded 13 tasks. Hill then wrote a second rule, rate, which tolerated a minority of disagreeing repetitions instead of discarding the whole task. On the same three repetitions, rate produced 13–2 over 15 informative tasks, with p=0.0074.

Hill acknowledges that he wrote rate after seeing the strict result. That sequence is the main reason to read the apparent improvement cautiously. The figures are his own, from one run of one suite, and he did not provide an independent reproduction.

The replication, and what pooling did

Hill then preregistered a replication on fresh repetitions 4 to 6. He predicted that strict would again fail to reach significance. It did not: it produced 9–1 over 10 informative tasks, with p=0.0215. Hill describes this as stronger than he predicted, not as a prediction he got right.

The title’s “twice the data, less power” refers to pooling the two three-repetition windows under strict. The table below lists the figures the article reports for each window.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Window and rule Repetitions Discarded tasks Informative tasks Paired result p-value
strict, first window 1–3 13 8 7–1 0.070
rate, first window 1–3 not stated 15 13–2 0.0074
strict, fresh replication 4–6 not stated 10 9–1 0.0215
strict, pooled 1–6 17 8 not stated 0.0703

The article frames the pooled change as discards rising while informative tasks fell, and p moving from 0.0215 to 0.0703. Its discard and informative counts for each window do not all line up in one place, so treat the table as the author’s per-window figures rather than a single reconciled ledger. The pooled paired result is not stated in the post.

Why more repetitions can cost retained tasks

Under strict, a task is removed if any repetition disagrees with the others. Each additional repetition gives another chance for a task to show a disagreement, so the set of tasks that survive can shrink as the run grows, even though the underlying questions have not changed. Hill’s interpretation is that the filter spends sample size: the more it observes, the more it has to discard.

He puts the trade-off this way:

“If the discard count rises faster than your informative count, your filter is spending your sample size, and the direction of that trade is not obvious from the code.”

He also writes: “A conservative rule is not a free choice.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Both lines are the author’s own. They describe his reading of his result and should not be taken as a general statistical standard.

Best Value
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the evidence does and does not establish

  • One suite and two models. The result comes from a single suite of trap questions. Hill expressly warns against reading it as a capability ranking of the models.
  • Timing is confounded with pooling. The first window and the replication were run at different times. Pooling them may mix sampling periods. Hill attributes the loss of power to the rise in discards, but he has not run the analysis that would separate that explanation from the timing difference.
  • rate may simply be the better rule. Hill concedes this, and concedes that the discard analysis could be a defence of a mistake. He treats the fresh-data replication as evidence against that reading, while noting that one replication with two models is limited.
  • Parameters. Hill states that rate_margin was fixed at 0.5 before the replication, and that the --alpha and margin flags were later removed from the command line. These are his accounts of implementation and preregistration; no external audit is cited.
  • Date. The article is marked “Sep 23” and does not show a year.

The post does not establish a universal optimal threshold, and it does not show that strict filters are worse in general. It shows that one strict filter, on one suite, lost retained tasks as its window grew.

How to check whether your filter is discarding too much

  1. Log two counts for every comparison: tasks discarded by the stability filter, and tasks retained as informative.
  2. Record both counts after each batch of repetitions you add, not only at the end of the run.
  3. If the discard count grows faster than the informative count, compute the paired result on the fresh batch alone as well as on the pooled data.
  4. If the fresh and pooled results disagree, check whether the batches were sampled at different times before you attribute the change to the filter.
  5. Write the filter’s thresholds and margins down before you look at the comparison, and report the rule you used alongside its result, including any rule written after an earlier result.

Implementation context in pi-eval

Hill links to the public pi-eval repository. Its README describes deterministic grading with fixed predicates and says it does not use an LLM judge for grading. It also documents suite fingerprinting and handling for inconclusive comparisons. Those features are useful context for the workflow, but they do not independently confirm the article’s findings.

Summary of the practical lesson

If you maintain an evaluation suite with a stability or flakiness filter, count what it discards and count again after you add data. A strict rule reduces the chance of counting unstable tasks as informative, but it also reduces the retained sample. Compare fresh and pooled windows when your schedule changes, and keep the rule’s parameters fixed before you interpret the comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The original account is Hill’s post on DEV Community.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.