Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A green test run tells you that the checks written into the suite passed under the conditions those checks create. It does not tell you the software will behave safely under failures nobody thought to encode. Derek Wang’s 2026 DEV Community essay, “A test system that can say ‘I don’t know’ is worth more than one that says ‘passed'”, turns that distinction into a practical design rule: a test system should separate failures it recognizes from failures it does not, and it should treat an unrecognized failure as evidence that the team’s map of failure modes is incomplete.

What a passing suite actually establishes

Every test encodes an expectation: given this input, in this environment, the system should produce that output. A pass confirms that the expectation held for the cases the test author chose. It says nothing about inputs that were never generated, dependencies that behaved differently from the test double, or conditions that exist in production but not in the test environment.

That is why a large test count can be reassuring without being decisive. The count measures how many expectations were written down, not how many ways the system can fail.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Known failures versus unmatched failures

Wang’s central move is to make the test system distinguish two outcomes that most suites collapse into one red or green signal:

  • A known failure matches an entry the team already understands. The shape of the failure, its root cause, and a repair recipe are on record, so the response is routine.
  • An unmatched failure matches nothing in the ledger. The essay treats it as an unknown and records it for investigation rather than filing it as one more red test.

The second category is the useful one. An unmatched failure is information about the team’s understanding, not just about the code. It means a failure mode exists that the team had not classified, which is exactly the situation a binary pass/fail report hides.

The failure ledger

The essay’s proposed mechanism is a ledger of failure portraits. Each portrait records three things: the known shape of the failure, its root cause, and a repair recipe. Each run is compared against this ledger before it is judged.

What goes into a portrait

  • Shape: the observable signature, such as which endpoint fails, what status code appears, and under what input pattern.
  • Root cause: the diagnosed reason, written once the team understands it.
  • Repair recipe: the steps that resolved it, so the next occurrence is handled the same way.

Where entries come from

According to the essay, entries can come from three sources: unmatched failures that are later investigated, issues that have been diagnosed and whose root cause is now known, and repair commits that fixed a real defect. Each source turns a surprise into a classification the suite can apply next time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Upkeep is organizational learning, not automation

Wang is explicit that the ledger does not maintain itself. Adding entries is a practice the team carries out between runs. A ledger that is never updated will classify fewer failures correctly over time, and a ledger that accumulates entries does not, by itself, make a system comprehensive.

Raising the baseline after improvement

The second practice is a regression baseline. Each run is compared with a recorded baseline. When results improve, the stronger state is recorded, so a later change cannot quietly drop back to the old level. When a regression appears, the useful output is not only “tests failed” but which failure class returned.

A baseline is only meaningful alongside the categories and scope it counts. A baseline that tracks a narrow set of failure classes can show stable numbers while other failures grow outside its view. Read it together with the ledger, not as a standalone measure of software quality.

The incident that passed every check

The essay’s most instructive example is a production failure that escaped a suite reported as fully green. A third-party endpoint returned a malformed response shape during a narrow time window. The error propagated through the chain of services that consumed it. The team reportedly could not reliably reproduce the anomaly, because it depended on a particular distribution of data that the test environment did not contain.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wang reads this as an unmodeled boundary rather than a simple shortage of tests. The suite was not missing a count of checks so much as missing a description of the condition that produced the failure. An integration check that exercised the same path with ordinary data passed, and it passed for the right reason: it tested what it was written to test.

The proposed response combines two things. Regression checks keep previously fixed failures from returning. Failure-mode analysis asks what happens when a failure cannot be reliably reproduced, and whether the system can tolerate it, contain it, or degrade safely. In the essay’s framing, resilience in parsing and degradation is the control for the cases the tests cannot pin down.

Reproducible testing and failure-mode analysis solve different problems

Reproducible tests answer whether a known behavior still holds. Failure-mode analysis answers a different question: when the system meets a condition nobody has reproduced, does it fail in a contained way? For the malformed-response incident, a reproducible test could only be written after the shape was understood. Failure-mode thinking asks earlier: what should the consumer do with a response that does not match the expected structure, and does that behavior stay bounded if the upstream service misbehaves again?

Figures from the author’s project

The essay reports several numbers from the author’s own project. They are useful as an illustration of the approach, but they are the author’s original data. They were not independently audited, and they are not benchmarks for other teams.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Figure What the author reports Qualification
10 seconds The full regression run completed in about ten seconds. Derek Wang, 2026, author-reported project figure. Depends on that project’s size and hardware, which the essay does not benchmark against other setups.
50 scenarios across six suites Coverage of contracts, idempotency, retrieval, regression, resilience, and end-to-end tests. Derek Wang, 2026, author-reported project figure. Scenario count reflects how the author defined scenarios.
22 hidden HTTP-500 errors The full suite caught all 22 before release. Derek Wang, 2026, author-reported. The essay does not provide independent confirmation of the count.
62/62 integration checks An integration test was fully green during a later production incident, and all six suites passed. Derek Wang, 2026, author-reported. This is the figure that illustrates the gap: a fully green result coexisted with a production failure.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why scoring rules reward guessing

A separate commentary, published on Claims · The Evaluating Self, explains a related incentive in answer scoring. Under a scoring rule that gives credit for a correct answer and zero for abstaining, replacing a calibrated “I don’t know” with a guess can raise the expected score. The commentary is careful to note that this describes the scoring rule and does not quantify how much the incentive explains real-world behavior.

The same logic is worth borrowing for test reporting, as an analogy only. If a suite reports pass counts and has no way to represent an unclassified result, every unknown is forced into a pass or a fail. Reporting “unmatched” as its own outcome is the test-system equivalent of permitting honest abstention. It is a different mechanism, but it addresses the same problem: a scoring structure that makes overconfident reporting the easiest path.

Checking whether your own suite can say “I don’t know”

The essay’s ideas translate into a short set of questions you can ask of any suite:

  • Does a result distinguish known failures, new failures, and unresolved outcomes, or is everything reported as pass or fail?
  • Is each failure class linked to a diagnosis and a repair, or are failures only counted?
  • Are baseline changes recorded explicitly, and does a regression name the failure class that returned?
  • Is there evidence that the suite caught earlier incidents, rather than only a raw number of tests?
  • For failures that are hard to reproduce, are there resilience controls such as validation, containment, or degraded behavior, and an incident path that feeds back into the ledger?

A suite that answers “no” to several of these may still be useful, but its green result should be read as a statement about the encoded checks, not about the system’s safety.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Limits of this evidence

The main source is a first-person engineering essay. Its time, suite-size, error-count, and integration figures are the author’s own project data. They have no independent confirmation, and they should be treated as one example of the approach rather than an industry study or a guaranteed outcome. The scoring-incentive commentary is a separate source about answer evaluation in AI systems; it is not evidence about software test suites, and the parallel drawn above is an analogy.

The essay does not claim that a ledger or a baseline makes software safe. Its argument is narrower: a system that can name what it does not recognize gives a team information that a green checkmark conceals.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.