Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To test AI-generated code independently, separate the implementation from the acceptance tests: give the coding agent the requirements, but keep the acceptance criteria from it while a separate tester writes tests from those criteria. This makes failures more informative, but a passing test is not proof that the code is correct—especially if the code already existed and its author may have known the criteria. The workflow is useful only when the requirements themselves state the intended behavior precisely and a domain expert approves them.

What independent testing changes

In a conventional workflow, a developer may implement a requirement and write tests with knowledge of the same expected outcomes. In a spec-driven workflow, the implementation and verification roles are separated: the coding agent receives the requirements, while a separate testing agent receives acceptance criteria and writes tests without seeing the implementation.

The point is not to make the test writer infallible. It is to reduce the chance that the verifier unconsciously reproduces the implementation’s assumptions. As Gal Arav puts the principle, “the person who builds the system must never be the person who verifies it.” This is the author’s formulation, not a quotation from an external standard. Arav’s September 30, 2026 article describes the method and its limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the information boundary real

Simply assigning different names or agents to the roles is not enough if both can inspect the same implementation and criteria. Preserve the separation in what each role can access: the coding agent should not receive the hidden acceptance criteria, and the tester should derive tests from the written criteria without inspecting the code. Record the inputs and outputs so the process can be reviewed later.

This reduces one source of bias; it does not eliminate other risks. A test writer can misunderstand a requirement, and a test suite can miss conditions that its fixtures never exercise.

What failures and passes establish

Interpret test results in light of who could see what and when. The strongest evidence in this workflow is a failure found by a tester who did not inspect the implementation and whose criteria were withheld from the coder. It shows a disagreement between the implementation and the stated bar. It does not, by itself, prove that the bar describes the right behavior.

Situation What a failing test indicates How to read a passing test
Code created during the separated workflow The implementation failed a criterion the coding agent had not seen. Investigate the implementation and the requirement. More meaningful evidence against criterion-specific gaming, but not proof of correctness or completeness.
Code that existed before the workflow A test failure remains a finding if the tester derived the test from criteria without reading the implementation. Weaker evidence: the original author may already have seen the criteria. Commit order can offer a limited clue, but commit dates do not establish when code was written or what its author knew.

A pass means the code satisfied the tests that ran. It does not establish that every relevant criterion was tested, that the tests cover realistic cases, or that the specification is valid.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make boundary behavior explicit before writing tests

Boundary cases expose underspecified requirements. Consider a warning for time headway below two seconds. Does exactly 2.00 seconds trigger the warning? “Below two seconds” ordinarily excludes equality, but wording such as “breaks the two-second rule” can invite different interpretations. Arav offers a useful diagnostic: “Given only the requirement, could two competent developers disagree about exactly 2.00 seconds?” If yes, settle the decision in the requirement rather than leaving it hidden in the acceptance tests.

Write the rule so both implementation and test are derivable

For example, specify: “Issue a warning when calculated time headway is strictly less than 2.00 seconds; do not issue one at exactly 2.00 seconds.” If the intended rule includes equality, say “less than or equal to 2.00 seconds” instead. Also define how invalid samples are handled, how headway is calculated, and what inputs count as invalid; the test writer should not have to infer these behavior-defining choices.

Arav reports that, in ten runs with an ambiguous first specification, three converged on the first sweep. After the boundary decision was added to the requirement, all ten reportedly converged on the first sweep, while seven still needed a repair for acceptance of a zero-metre gap. These are the author’s results for that example, not independently reproduced or general benchmarks. The lesson is to clarify the requirement, not to loosen a test simply to obtain a pass.

How to apply the workflow to an AI coding task

  1. Write the behavior contract. State valid and invalid inputs, calculations, outputs, error handling, and exact boundary behavior. Have a domain expert approve the intended behavior.
  2. Separate access. Give the coding agent the requirements it needs to implement the feature, but withhold the acceptance criteria. Give the test author the criteria and relevant interface or input contract, not the implementation.
  3. Build tests from criteria. Turn each criterion into observable assertions, including equality boundaries and invalid-input cases. Use fixtures that can actually trigger each rule.
  4. Run tests against the implementation. Treat failures as specific disagreements to investigate. Fix the implementation when it is wrong; revise the specification only when the intended behavior was genuinely unclear or incorrect, with expert approval.
  5. Review what the suite did not establish. Check for untested criteria, unreachable fixtures, missing edge cases, and validation questions about whether the written behavior matches the real need.

Test data must exercise the rules

A rule that no fixture can trigger is not meaningfully tested. If a system rejects zero-metre gaps, include a zero-gap sample and verify the expected result. If a warning threshold is 2.00 seconds, include values below, at, and above the threshold. For systems with rare but consequential scenarios, average performance can conceal failures on specific cases such as cut-ins or occlusions. The test set must represent the conditions the criteria are meant to govern.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Exhaustively enumerating edge cases is often impractical, particularly for advanced driver-assistance systems and their operational design domains. Arav points to design-of-experiments principles as a way to choose informative coverage rather than relying on brute force. Whatever sampling strategy is used, human judgment is still needed to decide which conditions matter.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Verification is not validation

Verification asks whether the implementation meets the written standard. Validation asks whether the standard describes the behavior that should actually happen. An independent testing agent can help with verification; it cannot decide whether the specification captures the right product, safety, or domain requirement.

Keep an accountable domain expert involved in approving and evolving the specification. A perfectly consistent implementation and test suite can still produce the wrong outcome if both follow a flawed requirement.

What the reported run does—and does not—show

Arav’s example concerned logged radar samples: reject invalid samples, calculate time headway, and warn below a two-second threshold. He reports that the coding agent initially accepted a zero-metre gap; a separate test based on criteria the coder had not seen exposed it, and the agent changed the lower-bound check. The author says the run took under a minute and fewer than ten model calls. Those timings and events are his account, not an independently reproduced test.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Across three sweeps, the article reports 967 runs: roughly eight in ten passed integration and system tests, and roughly six in ten passed all stages, including unit tests. A separate fourth sweep included 390 runs; Arav says it reproduced approximately the same rates after process hardening and making two tasks harder. He reports that sweep separately rather than pooling it with the first three.

The author describes the model as small and inexpensive and frames the results as a performance floor. The figures describe his example tasks and process; they do not establish that hidden criteria catch more real defects than tests written with full code access, or that automatically refining criteria improves test quality. Arav identifies both as open questions requiring formal proof.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.