Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Generated unit tests can pass while confirming a bug. That happens when a test copies what the code does instead of checking what the requirement says it should do. A green test run is useful evidence, but it cannot by itself prove the behavior is correct.

How a test can pass and still be wrong

In his DEV Community article, originally published at TestingIL, Gil Zilberfeld describes this as a “tautological test”: the expected result reflects the implementation’s behavior, so the test passes even though the implementation violates the requirement. The test may look plausible and run successfully, yet provide false confidence.

The key question is not only whether the test passes. It is whether its expected result follows from the stated requirement independently of the code being tested.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The expiration-date example

Zilberfeld illustrates the problem with a function that checks a credit card’s expiration date. The function compares the first day of the expiration month with the current date. As a result, it can report a card expired before that month has ended.

In the example, a generated test for a card expiring in the current month expects True in the middle of that month. Zilberfeld argues the expected result should be False: the card remains valid through the end of its expiration month. The test passes against the function’s behavior, but the expected result conflicts with the rule the example is meant to express.

This is an illustrative example from the article, not a measurement of how often generated tests fail this way.

Why did it generate the wrong test?

A test generator can infer likely cases from the code it sees. If the implementation already contains a mistaken assumption, a generated test may encode that same assumption in its expected result. The test then checks consistency with the implementation rather than correctness against the requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This is why a test suite can grow without providing a matching increase in confidence. More passing tests do not necessarily mean more independently verified behavior. As Zilberfeld puts it, “The irony is that while we finally got more tests, our confidence in them is lower.”

How to review generated tests

Review tests as claims about required behavior, not just as code that must turn green. Zilberfeld recommends making test provenance and human review explicit, then paying particular attention to risky logic and tests that pass when they should fail.

  • Identify generated tests. Ask which tests were generated and which are trusted. Keep their origin visible so reviewers know where expected results came from.
  • Check the requirement first. For each important test, compare its expected result with the stated rule, not merely with the current implementation.
  • Review edge cases in risky logic. Test boundaries where a small difference changes the outcome, such as dates near the start or end of an expiration month.
  • Look for tests that should fail. Ask whether a test would catch a plausible violation of the requirement. If incorrect behavior still passes, revise the expected result or add a stronger case.
  • Clarify what received human review. A reviewer should assess both the code and the test expectations; reviewing only the implementation leaves the test’s assumptions unchecked.
  • Teach the practice across the team. Make provenance, requirement-based expectations, and edge-case review part of normal test development.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What a green result does—and does not—establish

A passing test establishes that the implementation produced the result the test expected for that input. To treat it as evidence of correctness, you also need reason to trust that expectation: it should come from the requirement or another independently justified source, rather than being inferred from the same potentially faulty implementation.

Zilberfeld’s article makes this practical argument but does not quantify how common the failure mode is or compare generated tests with human-written tests. Its lesson is narrower and useful: test counts and green runs are not substitutes for checking whether expected results match intended behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read Gil Zilberfeld’s article on DEV Community.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.