Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Passing tests show that code ran without triggering a failure; they do not prove the tests would catch a defect. In Marvin Okafor’s 2026 example, an AI model generated 69 tests for a Python module. Every test passed on the original code, but none detected 11 deliberately planted bugs. A larger experiment using mutation testing found that targeted test generation caught more faults in selected modules—but the results are specific to that experiment, not a general verdict on AI-written tests.

Why passing tests can miss bugs

A test can execute a function and still fail to check whether the function returned the right value, raised the right exception, or handled an edge case correctly. Such a test may pass on both correct and defective code. That is why a green test run, by itself, is not evidence that a suite can distinguish the intended behavior from a bug.

In Okafor’s initial example, the 69 generated tests all passed against the clean Python module, yet none caught its 11 deliberately introduced bugs. This is an illustration from one module, separate from the author’s later comparison across twelve library targets. It shows the difference between tests that run and tests that detect faults; it does not establish how often AI-generated tests fail this way in other projects.

What mutation testing measures

Mutation testing makes small, deliberate changes to source code—such as flipping a comparison, changing a constant, or removing a raise—and runs the test suite against each altered version. Each change is a “mutant.” If the suite fails on a mutant, it detected that change; if the suite still passes, the mutant survived and represents a possible fault the tests did not catch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Line coverage answers whether a line ran during a test. Mutation testing asks whether a selected change to code would make the test fail. Neither metric alone proves that a test suite is complete or that it catches every real-world defect: mutation results depend on which changes are tried, which code is reachable, and how the experiment is measured.

What the twelve-library experiment found

In the larger experiment, Okafor reports generating 455 mutations across twelve Python-library targets. Existing test suites let 133 mutations survive, but only 53 of those were on lines the suites actually executed. The other 80 surviving mutations were on unexecuted lines, so those tests could not assess their behavior.

Among the 53 reachable surviving mutations, Okafor compared three test-generation approaches. The reported results are the author’s findings, not independently replicated estimates:

Approach Mutation hint Generation and acceptance rule Reported detections
Targeted generation with a pass/fail gate The model received a specific mutation hint. Keep a generated test only if it passed on clean code and failed on the targeted mutant. The author says the gate used a subprocess exit code rather than a model’s judgment. 44 of 53 reachable surviving mutations.
One broad “write more tests” prompt No specific mutation hint. One broad prompt; the article says the approaches used the same model and token ceiling. 9 of 53 reachable surviving mutations.
One untargeted test per call No specific mutation hint. One untargeted test per call; the article says the approaches used the same model and token ceiling. 2 of 53 reachable surviving mutations.

The denominator matters: those counts concern the 53 mutations that survived existing tests on code those suites executed, not all 455 mutations or all possible defects. The repository’s description further narrows the question to whether a mutation hint, execution gate, and one-test-per-call setup outperform comparison conditions over reachable survivors in selected modules. It explicitly says the experiment is not a general measure of whether agents write good tests. See the killcheck repository and Okafor’s article.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Did the targeted tests generalize?

The initial comparison evaluates whether generated tests caught the mutations they targeted. It does not by itself show that those tests would catch different mutations or faults elsewhere. Okafor reports that 44 retained targeted tests had zero reported cross-function transfer, while 36 caught exactly one mutation. The later repository update adds an important qualification: transfer was observed within functions, but not across functions.

In that update, a frozen set of 44 tests caught 34 of 53 fresh reachable mutants. The author also reports that the pooled fresh population was 92—below a preregistered minimum of 100—and that two targets accounted for 30 of the 53 reachable mutants. These constraints make the holdout result informative but limited; it should not be presented as broad proof that targeted tests reliably generalize.

What the experiment says about coverage and test quality

For these selected targets, the difference between 133 surviving mutations and 53 survivors on executed lines suggests that unexecuted code was a substantial part of the gap. Okafor also says widening the test commands by six to forty times changed the reachable-survivor count from 54 to 53. That finding supports the author’s interpretation for this experiment, not a universal claim that weak assertions matter less than missing coverage in every codebase.

Coverage and mutation testing answer different questions. Coverage can help locate code that tests never reach. Mutation testing can expose tests that reach code but do not fail when its behavior changes in a selected way. A useful assessment therefore treats them as complementary signals, rather than assuming that high coverage guarantees meaningful assertions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why the harness matters

The reported results depend on the toolchain that creates mutations, runs tests, and classifies outcomes. Okafor says the harness had 11 bugs, followed by three more issues readers found after publication. The author reports problems including editable installs hiding mutations, parallel execution corrupting a target, a classifier using the wrong unit, stale bytecode, and a pytest outcome bucket matching a string the installed version did not emit.

According to Okafor, each problem either made results look better or made an absence of evidence look like evidence. The author says predicted-outcome checks exposed instrument problems that code review alone had not found; readers later identified further issues by examining those checks. This is a project-specific debugging account, not evidence that every software evaluation is biased. It does show why a mutation-testing result is only as trustworthy as the harness and outcome rules behind it.

How to use the lesson in a real test suite

For a developer deciding whether tests are worth keeping, ask what behavior each test protects—not just whether it executes code. A focused test should pass when the implementation is correct and fail when a relevant behavior is broken. Mutation testing can help probe that distinction, while coverage can point to code that has not been exercised.

  • Use mutations as probes, not as a complete catalog of possible bugs; a surviving mutant identifies a gap for investigation, not necessarily a real defect.
  • Keep the clean-code check: a generated test that fails before any mutation is applied is not evidence that it detected the intended fault.
  • Track the denominator and reachability. Report separately which mutations were generated, which survived existing tests, and which were on executed lines.
  • Validate the harness with cases whose expected outcomes are known, and make the test commands, classification rules, and target scope inspectable.
  • When evaluating generated tests, distinguish catching the exact mutation shown to the model from transfer to new mutations, functions, or projects.

Okafor’s practical recommendation is to publish the evaluation harness when building evaluations for one’s own work. The killcheck repository’s scope warning is equally important: its results concern a bounded measurement setup, not the overall quality of AI agents’ tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.