Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Marvin Okafor built an agent to strengthen weak tests, then found eight defects in the harness measuring whether it worked. In his account, each defect could make results look better, cleaner, or more publishable than they were. The lesson is not that the agent proved a general improvement in software quality; the experiment reached only a small, selected sample. It is that an evaluation can mislead even when its numbers look precise—and that the instrument deserves the same scrutiny as the system being measured.
What the agent tried to fix
Line coverage tells you that tests executed code; it does not tell you whether they would notice incorrect behavior. Mutation testing probes that gap by making small changes to a program and checking whether the test suite fails. A surviving mutant signals that the tested suite did not detect that particular change. It can reveal a testing gap, but does not by itself prove the original program has a production defect.
Okafor’s agent focused on surviving mutants. It sent a mutation diff to a model and asked it to write a test. The harness accepted a generated test only when it passed against clean code and failed against the targeted mutant. On a retry, the model could receive actual pytest output. The acceptance decision relied on the subprocess exit code, not on the model’s own claim that its test worked.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →What the reported experiment found
In the author’s account, the project examined 455 mutants across 12 Python libraries; 133 survived the existing suites. Of those survivors, 53 were on lines the tests executed. That contrast illustrates why coverage and fault detection answer different questions: execution alone does not establish that assertions would catch a wrong result.
Okafor says that widening per-target test commands by between 6 and 40 times changed the count from 54 to 53, as previously unreachable mutations became kills. These are figures from his project, not independent or industry-wide benchmarks.
The reported comparison covered 15 mutants: a single-test baseline killed 1, while the agent killed 9, with a 60% keep rate. The author says the agent ran on only 2 of 10 targets before its API budget ran out, and that these were the targets where the baseline did worst. The sample was therefore neither random nor representative. All nine kills were on the two cheapest mutation types; the author reported no cross-function transfer, and seven kept tests killed only their targeted mutation.
The outcome also differed from Okafor’s expectation. He anticipated that accepted tests might often be vacuous—tests without meaningful assertions that happened to fail on a mutant. Instead, he reports an empty “none” category, eight of nine kills as real assertion failures, and six discarded drafts that passed on clean code but did not detect the mutation. In this small run, the gate filtered tests that were valid but ineffective, rather than simply rejecting broken tests.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThe eight defects in the measuring harness
Okafor says each of the following defects would have made the results look better, cleaner, or more publishable. That direction-of-bias conclusion is his account; the underlying code is not available here for independent audit.
- Mutations were applied to a copy the tests did not import. For src-layout packages, editable installs resolved imports to the original checkout while mutations were written to a temporary copy. Mutations were therefore invisible, making three targets appear to score 0.000.
- Parallel execution made one target unstable. Concurrent mutant runs produced three different survivor sets across four runs for a target using asynchronous I/O.
- The test-file picker chose the wrong file. On the hardest target, the harness selected a test file that was not the relevant one.
- A batch-level classifier overstated individual test strength. It classified batches rather than individual tests, so one strong test could make an entire batch of 69 look strong.
- Reconstruction dropped shared imports. That step manufactured failures by omitting imports used across the tests.
- The extractor discarded valid unittest responses. It scanned only top-level functions, dropping valid
unittest.TestCaseresponses and sending a harness error—not pytest output—to the retry loop. - An assertion check misclassified unittest syntax. It treated
self.assertEqual(...)as “no assertion,” potentially creating the very outcome the author had hypothesized. - A metric was undefined for some Python methods. A pre-registered metric did not apply to dunder-dispatched code such as
__call__and__or__, yet appeared as a real, near-zero rate.
How the defects were discovered
Okafor describes a practical way to expose errors that ordinary code review had missed: predict what a check should report before running it, then investigate when the observed result differs. As he put it, “None of them was found by reading code.” Each, he says, surfaced when he predicted a check’s outcome in advance and found that the outcome was wrong.
His question is useful beyond mutation testing: “What would my instrument look like if it were lying to me?” A check should have an expected result, and its likely error should be considered in terms of direction: would it inflate a success rate, hide a failure, or make a metric look more precise than it is?
Rank #4
What reproducibility could—and could not—show
For a clean-clone check, Okafor says he ran each target three times serially and required byte-identical survivor sets. Eleven of the twelve targets reportedly reproduced the same set across those runs; one varied, and the result was disclosed in the README. Serial repetition made instability visible, but does not establish that the remaining targets or the agent’s effect were independently reproduced.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsThe account describes a project built in about 30 hours for a challenge with around 7,800 registrants. It also says the author missed the submission deadline by 11 minutes. Those details provide context, not evidence of quality or performance.
Best Value
What readers can reasonably conclude
The strongest conclusion is methodological: a test-generation result is only as trustworthy as the path from code mutation to test execution, extraction, classification, and reported metric. A gate that verifies both clean-code success and mutant failure is more informative than accepting a test because it runs or adds coverage. But the author’s small, selectively covered experiment does not establish that the agent generally improves test quality.
Okafor’s compact advice captures the point: “Before you measure an agent, write down what your instrument would look like if it were lying to you.”
Read Marvin Okafor’s original DEV Community article.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

