What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A strict stability filter can lose statistical power as you add repetitions, even when the models under comparison have not changed. In Erik Hill’s account of a 159-task evaluation suite, a rule that discarded any task whose configuration disagreed with itself produced a significant paired result on three repetitions, then a weaker result once the discard count grew and the informative count shrank after more data were pooled. The lesson is about the filter and how its retained sample changes, not about which model is better.
What the strict rule does
Hill compared two models on a frozen suite of 159 tasks. Each task was run across repeated trials, and the suite is built from trap questions. His first rule, strict, discarded a task if a single configuration disagreed with itself across repetitions. Only the tasks that survived were treated as informative and fed into a paired comparison between the two models.
The filter is attractive for a simple reason: a task whose answer flips between runs is noise, and counting it as a win or loss inflates confidence. The trade-off is that the filter also removes data, and the amount it removes depends on how many repetitions you run.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →The first result, and the rule written after it
Across repetitions 1 to 3, strict produced a 7–1 comparison over 8 informative tasks, with p=0.070. The filter had discarded 13 tasks. Hill then wrote a second rule, rate, which tolerated a minority of disagreeing repetitions instead of discarding the whole task. On the same three repetitions, rate produced 13–2 over 15 informative tasks, with p=0.0074.
#1 Best Overall
Hill acknowledges that he wrote rate after seeing the strict result. That sequence is the main reason to read the apparent improvement cautiously. The figures are his own, from one run of one suite, and he did not provide an independent reproduction.
The replication, and what pooling did
Hill then preregistered a replication on fresh repetitions 4 to 6. He predicted that strict would again fail to reach significance. It did not: it produced 9–1 over 10 informative tasks, with p=0.0215. Hill describes this as stronger than he predicted, not as a prediction he got right.
Rank #2
The title’s “twice the data, less power” refers to pooling the two three-repetition windows under strict. The table below lists the figures the article reports for each window.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match| Window and rule | Repetitions | Discarded tasks | Informative tasks | Paired result | p-value |
|---|---|---|---|---|---|
strict, first window |
1–3 | 13 | 8 | 7–1 | 0.070 |
rate, first window |
1–3 | not stated | 15 | 13–2 | 0.0074 |
strict, fresh replication |
4–6 | not stated | 10 | 9–1 | 0.0215 |
strict, pooled |
1–6 | 17 | 8 | not stated | 0.0703 |
The article frames the pooled change as discards rising while informative tasks fell, and p moving from 0.0215 to 0.0703. Its discard and informative counts for each window do not all line up in one place, so treat the table as the author’s per-window figures rather than a single reconciled ledger. The pooled paired result is not stated in the post.
Why more repetitions can cost retained tasks
Under strict, a task is removed if any repetition disagrees with the others. Each additional repetition gives another chance for a task to show a disagreement, so the set of tasks that survive can shrink as the run grows, even though the underlying questions have not changed. Hill’s interpretation is that the filter spends sample size: the more it observes, the more it has to discard.
He puts the trade-off this way:
“If the discard count rises faster than your informative count, your filter is spending your sample size, and the direction of that trade is not obvious from the code.”
He also writes: “A conservative rule is not a free choice.”
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Both lines are the author’s own. They describe his reading of his result and should not be taken as a general statistical standard.
Best Value
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
What the evidence does and does not establish
- One suite and two models. The result comes from a single suite of trap questions. Hill expressly warns against reading it as a capability ranking of the models.
- Timing is confounded with pooling. The first window and the replication were run at different times. Pooling them may mix sampling periods. Hill attributes the loss of power to the rise in discards, but he has not run the analysis that would separate that explanation from the timing difference.
ratemay simply be the better rule. Hill concedes this, and concedes that the discard analysis could be a defence of a mistake. He treats the fresh-data replication as evidence against that reading, while noting that one replication with two models is limited.- Parameters. Hill states that
rate_marginwas fixed at 0.5 before the replication, and that the--alphaand margin flags were later removed from the command line. These are his accounts of implementation and preregistration; no external audit is cited. - Date. The article is marked “Sep 23” and does not show a year.
The post does not establish a universal optimal threshold, and it does not show that strict filters are worse in general. It shows that one strict filter, on one suite, lost retained tasks as its window grew.
How to check whether your filter is discarding too much
- Log two counts for every comparison: tasks discarded by the stability filter, and tasks retained as informative.
- Record both counts after each batch of repetitions you add, not only at the end of the run.
- If the discard count grows faster than the informative count, compute the paired result on the fresh batch alone as well as on the pooled data.
- If the fresh and pooled results disagree, check whether the batches were sampled at different times before you attribute the change to the filter.
- Write the filter’s thresholds and margins down before you look at the comparison, and report the rule you used alongside its result, including any rule written after an earlier result.
Implementation context in pi-eval
Hill links to the public pi-eval repository. Its README describes deterministic grading with fixed predicates and says it does not use an LLM judge for grading. It also documents suite fingerprinting and handling for inconclusive comparisons. Those features are useful context for the workflow, but they do not independently confirm the article’s findings.
Summary of the practical lesson
If you maintain an evaluation suite with a stability or flakiness filter, count what it discards and count again after you add data. A strict rule reduces the chance of counting unstable tasks as informative, but it also reduces the retained sample. Compare fresh and pooled windows when your schedule changes, and keep the rule’s parameters fixed before you interpret the comparison.
The original account is Hill’s post on DEV Community.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

