Most of the judged changes in a 2026 study had at least one edited or added test that failed when the change’s source code was reverted. But a passing test run alone cannot show that a test detects a fix: to check that, the study ran the same tests against the old source. It found that 64 of 71 judged maintainer fixes (90%) and 75 of 91 judged agent-authored pull requests (82%) were proven by at least one test. Those figures describe selected samples, not software development or coding agents as a whole.
What does it mean for a test to notice a fix?
A regression test is useful evidence for a change when it passes with the change and fails without it. Receipts, an open-source tool, tests this counterfactual by running a change’s added or edited tests twice: first with the change applied, then after restoring only the changed source files to their parent-commit or pull-request merge-base versions. The tests, dependencies, and configuration stay at their newer versions in both runs.
This is a focused check: does a test react to the code change? A test that fails against old source is evidence that it notices a difference, not proof that the change is correct, that it covers every relevant behavior, or that the original bug is fully fixed. [Receipts project]
What did the 181-change study find?
The study, published by syntaxixr in 2026, examined 81 maintainer fix commits and 100 agent-fingerprinted pull requests across 17 open-source projects. Ten maintainer changes and nine agent pull requests could not be judged for environment reasons, so percentages use the remaining judged changes.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems| Sample | Selection | Judged | Proven | Other reported result |
|---|---|---|---|---|
| Maintainer fixes | 81 commits from 12 libraries | 71 of 81 | 64 of 71 (90%) | 10 were not judged; the study’s summary does not report a weak-only figure for this sample. |
| Agent-fingerprinted pull requests | 100 newest qualifying PRs from five agent-heavy repositories; open or closed, including unmerged work | 91 of 100 | 75 of 91 (82%) | 9 of 91 (10%) were weak-only. |
These are the study’s counts and percentages, not independent estimates of how often tests or AI-generated code are effective. The samples were small and non-random. Maintainer commits were selected from well-maintained libraries; the agent sample came from five repositories selected for agent-heavy activity. The study explicitly limits its conclusions to the projects it examined. [Study method and results]
What do “proven,” “weak-only,” and the other verdicts mean?
Receipts assigns test-level verdicts, then summarizes them at the change level. In this study, a change counted as proven only if at least one test failed against old source and none of its tests were weak or theater.
- Proven: At least one test fails without the change, and none are weak or theater.
- Mixed: Some tests prove the change, while others are weak.
- Unproven: Every tested change-related test also passes against old source.
- Weak-only: Tests fail against old source only because they call code that did not exist yet.
- Guard: A test passes on both sides but sits alongside a test that proves the change.
- Theater: Tests pass both with and without the change, and no test proves it.
- Broken, flaky, or skipped: Separate outcomes for tests that fail for other reasons, behave inconsistently, or do not run.
“Unproven” does not necessarily mean that no test is useful; it means the selected tests did not demonstrate a difference under this procedure. The categories also help distinguish a genuine test failure on old behavior from a test that merely cannot run in the old environment.
Why can a test fail against old code without proving the fix?
In nine judged agent-authored pull requests, the study classified the evidence as weak-only. In the pattern it describes, a test module imports a name added by the change at the top level. When Receipts restores old source, that name is missing, so the test module fails to load before the tests can exercise the old behavior. The failure looks like a signal, but it does not show that the test would catch the original regression.
Recommended Free Tools
The article illustrates this with a Claude Agent SDK Python pull request and suggests importing newly introduced names inside only the tests that need them. That lets unrelated tests load against old source and makes the counterfactual more informative. [Study article and example]
What makes the comparison between maintainer and agent changes difficult?
The agent sample was attributed by fingerprints, not verified authorship. Of its 100 pull requests, 87 carried Claude Code fingerprints; the study also counted six Codex, six Cursor, and one Copilot fingerprint. A fingerprint is evidence about the tooling associated with a PR, not a guarantee that an agent acted without human steering. The study notes that human involvement is possible.
Rank #4
The samples also differ in how they were assembled: up to eight recent qualifying maintainer commits per repository were drawn from 12 libraries, while up to 20 newest agent-fingerprinted PRs per repository were drawn from five agent-heavy projects. Agent PRs could be open or closed and included unmerged work. The different selection methods and repositories mean that the 90% and 82% figures should not be read as a controlled head-to-head test of maintainers versus agents.
Why might the test verdict miss a real change?
The study says theater was uncommon, but gives examples showing why test outcomes depend on the change and environment. A type-only change may not affect runtime tests. A Windows-specific newline fix may not be observable when tests run on Linux. A dateutil representation change may produce output matching inherited behavior. A maintenance commit may mention an issue without making a behavior change that its tests can detect.
Best Value
These cases do not make counterfactual testing pointless; they define its limits. A test suite can be unable to demonstrate a particular change even when the change has a valid purpose. Conversely, a failure against old source can arise from an import or environment problem rather than from the behavior the fix targets.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How can a team use Receipts?
The Receipts README describes the project as deterministic, with no LLM or API key requirement, and says it runs the project’s own test runner. It documents support for pytest, vitest, and jest, along with a command-line interface, GitHub Action, and an agent skill for tools that read SKILL.md. Its stated baseline is Node 20 or later and Git. These are documented project capabilities, not independent test results. [Receipts README]
Run it locally or in a pull request
The project provides reproduction commands and raw results alongside its study materials; consult the README and study for the exact commands and repository-specific setup rather than assuming every project uses the same test configuration. For automated checks, the README documents a GitHub Action that can report on pull requests and fail checks for configured verdicts. Its example recommends triggering on pull_request and setting checkout credentials not to persist. A comment report requires comment permission.
The study also points readers to a Hugging Face dataset and repository results as its reproduction path. The published findings should be treated as the authors’ reported results: the complete study was not independently rerun for this article. [Study reproduction materials]
Free tools Windows power users keep installed
One-click scans. No signup required.
What should readers conclude from the results?
The study offers evidence that many changed tests in these selected projects responded to their fixes when the source was reverted. It also identifies a concrete weakness worth checking for: tests that import newly added names at module load time can fail against old source without exercising old behavior. The key distinction is between a test that notices a code difference and proof that the code change is right. The sample results illuminate that distinction; they do not establish a general success rate for coding agents or regression tests.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

