Use unit tests to check isolated logic and integration tests to check whether connected parts work together. For AI-generated code, neither kind of test is trustworthy just because a model wrote it: review each test against agreed requirements, run it in the project’s real environment, and inspect what it actually exercises.
What unit and integration tests each tell you
Unit and integration tests answer different questions. A layered test suite uses each where it can provide useful evidence; one is not a substitute for the other. ISO’s overview of AI-system testing describes test levels including unit/component, integration, system, system integration, and acceptance testing. Teams may define boundaries differently, so use the project’s own terminology consistently. ISO/IEC TS 42119-2:2025
| Aspect | Unit/component test | Integration test |
|---|---|---|
| Question | Does this isolated function or component behave as required? | Do connected components or services work together across the boundary being tested? |
| Dependencies | Usually replaces external dependencies with controlled mocks or stubs when those dependencies are not the subject of the test. | Exercises the interaction under evaluation, using real or representative dependencies as feasible. |
| Setup and feedback | Usually fast and isolated; useful for deterministic logic and frequent runs. | Often needs more configuration and can reveal boundary, contract, data-flow, or configuration problems. |
| Value for AI-generated code | Can catch local logic errors, boundary inputs, error handling, and transformation mistakes. | Can reveal incompatibilities or coordination failures that isolated tests cannot expose. |
| Main limitation | May assert the wrong behavior or mock away the defect. | Environment and service variability can make tests slower or less stable; keep scope intentional. |
This division follows established test-level guidance and recommendations to isolate dependencies where appropriate while using layered testing. ISO AWS Prescriptive Guidance AWS guidance on the testing pyramid
When to write a unit test—and when to add an integration test
Choose a unit test for isolated, deterministic behavior
Use a unit or component test when the behavior can be checked in isolation and the expected result is clear. For example, test a function that validates an input, transforms a response, or chooses an error path without making a live network request. If the component calls an LLM or another external service, a mock or stub can supply a controlled response so the test checks how your code handles it.
A mock is useful only if it leaves the behavior you care about intact. If the test mocks the very logic or interaction where a defect could occur, a pass says little about that risk.
Choose an integration test for an important boundary
Add an integration test when the interaction among components, APIs, tools, or workflow steps is itself important. This is where you can check, for example, whether one component sends the expected data to another and handles the response according to the agreed contract. Use controlled or representative dependencies where practical, and reserve live-service checks for cases where the real interaction is what needs evaluation.
For agentic systems, isolated exact-match unit tests may miss behavioral failures across prompts, tools, and workflow steps. AWS recommends testing across broader layers for these distributed systems. AWS Prescriptive Guidance for testing agentic AI systems
How to review and run AI-generated tests
Treat AI-written tests as proposed code, not as independent proof that generated application code is correct. A test can pass while encoding an incorrect assumption, mirroring the implementation, or failing to exercise the intended behavior. Microsoft’s VS Code guidance notes that adding tests to an existing project involves more than generating test code. Test existing code with AI
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Establish the project’s constraints. Read the relevant requirements and acceptance criteria. Identify the test framework, existing fixtures and helpers, project conventions, and the command used to run tests.
- Ask for proposed cases before test code. Include normal behavior, boundary values on both sides, invalid inputs, and relevant error cases. Where a requirement is unspecified, decide what the intended behavior should be rather than letting the model invent it.
- Agree on the cases and expected results. Review each proposal against observable behavior and requirements. Then ask for test-only changes, explicit expected values, and reuse of established helpers.
- Pick the layer based on the risk. Keep deterministic component checks isolated; use mocks or stubs for external services when testing surrounding logic. Add integration coverage for interactions whose contracts or coordination matter.
- Run the project’s actual test command. Inspect failures, skipped tests, and warnings instead of relying on a tool’s summary. Confirm that the intended code ran and that mocks did not replace the behavior the test claims to check.
- Use coverage as a map, not a verdict. Coverage can point to code that lacks tests, but does not show whether assertions capture requirements. Mutation testing offers a stronger check of whether assertions detect deliberately introduced faults.
- Run deterministic checks in CI. Automated checks on changes to deterministic application logic provide repeatable feedback as code evolves.
Why AI-generated tests need a human oracle
Testers need a reliable way to decide what result is correct. ISO/IEC TR 29119-11:2020 calls this the test-oracle problem: “testers find it difficult to determine expected results for testing and therefore whether tests have passed or failed.” Its guidance concerns testing AI-based systems generally, including black-box approaches and neural-network-specific white-box testing; it should not be confused with testing ordinary software merely because a code-generation model authored it. The ISO page lists the 2020 document as published and under review. ISO/IEC TR 29119-11:2020
For software that calls a nondeterministic AI service, unit tests can still check deterministic surrounding behavior with controlled responses. Testing actual service interactions or overall quality requires integration or system-level evaluation with criteria suited to the application. A single exact output is not necessarily an appropriate success condition for a variable AI response.
Rank #4
What published AI test-generation results do—and do not—show
Benchmarks are useful evidence about their stated setup, not a general warranty for generated tests. TestGenEval’s ICLR 2025 paper covers 68,647 tests across 1,210 unique code-test file pairs. In its evaluated setup, GPT-4o averaged 35.2% coverage and an 18.8% mutation score. Those figures describe that historical benchmark, not current model rankings or the expected quality of tests in a particular project. The paper’s use of both coverage and mutation score also illustrates why coverage alone is not a correctness measure. TestGenEval paper
NIST’s 2025 GenAI (Pilot) Code Challenge evaluates generated unit tests for elementary Python code. Its scope is a pilot: it does not establish performance across languages, large repositories, integration tests, or production systems. NIST GenAI (Pilot) Code Challenge
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsQuick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

