Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Good engineering is not measured by how few unit tests a team writes. It is measured by whether its tests catch the failures that matter, quickly enough to help developers fix them. Mock-heavy unit tests can miss defects at real dependency boundaries, but unit tests remain valuable for isolated logic. A resilient strategy uses unit, integration, and end-to-end tests for the different questions each can answer.

What the title gets right—and wrong

Tarek Mostafa’s article, “The Best Engineers I Know Don’t Write Unit Tests”, makes a pointed argument against tests that check interactions with mocks while failing to exercise real dependencies. That is a useful warning about test design, not evidence that experienced engineers avoid unit tests. The article itself recommends unit tests for isolated algorithms such as cryptographic functions, parsers, and mathematical logic.

Mostafa also recounts a payment-service incident and cites 94% code coverage and a 30% engineering-time cost. Those figures are claims in his article, not independently established findings in the sources available here; they should not be generalized into evidence about teams overall.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What each test layer tells you

Google’s guidance distinguishes the layers by what they exercise: unit tests cover a functional unit, typically with external dependencies mocked or faked; integration tests exercise a small group of units together; and end-to-end tests check a product journey as a user would experience it. Passing one layer does not prove the behavior covered by another.

Test layer What it exercises What it is useful for Key limitation
Unit A functional unit, often with external dependencies replaced by mocks or fakes Fast feedback on isolated logic and relatively local failure diagnosis A simulated collaborator may not behave like the real dependency
Integration A small group of components working together Checking interactions and boundaries that isolated tests do not exercise Usually involves more setup and can give slower feedback than an isolated test
End-to-end A product journey exercised as a user would exercise it Checking critical user-facing flows across the system It covers broader behavior, but does not replace focused tests at smaller layers

Google’s descriptions and recommendations are set out in its 2021 overview of testing at Google and its 2024 discussion of flaky tests and the test pyramid. The practical choice involves trade-offs: isolation can improve speed and help localize failures, while real or representative collaborators can reveal mismatches a mock cannot.

Where mock-heavy tests can mislead

A mock can confirm that code made an expected call to a simulated object. It cannot, by itself, establish that a real service, database, message broker, or payment provider will accept the call or behave as the test assumes. If the mock and the real dependency diverge, a test may pass while the integration fails.

This does not make mocks inherently bad. They can keep unit tests fast and controlled, especially when the goal is to test a unit’s decision-making in isolation. The risk is treating that result as proof of a boundary that the test never crossed. Use integration or contract checks where compatibility and interaction behavior matter; reserve end-to-end coverage for the journeys whose failure would most directly affect users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to choose a useful mix

  1. Start with the failure you need to catch. For a deterministic parser or mathematical rule, a focused unit test may provide the clearest feedback. For a database query, service interaction, or dependency contract, include a test that exercises that boundary.
  2. Match the test to the claim. A unit test can show that isolated logic behaves as expected under defined inputs. An integration test can check that connected components work together. An end-to-end test can check a critical user journey. Avoid claiming more than the test actually exercises.
  3. Consider feedback and upkeep. Favor fast, focused tests when they give adequate confidence. Add broader tests where the extra fidelity justifies their setup, maintenance, and possible reliability costs.
  4. Cover high-impact journeys deliberately. Select end-to-end tests for critical flows rather than trying to make them the only evidence of correctness. Add performance, load, or fault-tolerance testing when those risks are relevant to the product.
  5. Review failures as evidence about the strategy. If production defects repeatedly occur at a boundary not exercised by the suite, add a test at that boundary. If tests are hard to diagnose or routinely fail for unrelated reasons, reassess their scope and setup.

Google’s 2024 article presents the test pyramid as a heuristic: in general, a suite has more unit tests than integration tests and more integration tests than end-to-end tests. It is a way to reason about speed and fidelity, not a universal quota. An older 2015 Google Testing Blog post offers 70% unit, 20% integration, and 10% end-to-end as a first-guess rule of thumb, while noting that the exact mix differs by team. Neither the pyramid nor those percentages guarantee quality.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why coverage percentage is not a verdict

Code coverage can show which parts of a codebase a test suite executed; it does not by itself show whether the tests checked the important behavior or crossed the necessary boundaries. A high percentage can coexist with tests that assert little, while a smaller but well-targeted suite may exercise critical risks. Choose targets in light of the software’s purpose, audience, and failure costs, rather than treating a single percentage as proof of quality.

As George Pirocanac put it in the 2021 Google Testing Blog guidance, “A lot depends on the type of software, its purpose, and its target audience.” That context also determines how much confidence should come from tests, operational monitoring, and other safeguards; production observability can help surface failures, but it does not replace pre-release checks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.