Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose each test layer by the failure it needs to catch: use focused tests for local behavior, integration or service tests for boundaries between components, and a small set of end-to-end tests for important risks that depend on the assembled system. The test pyramid is a useful starting point, not a universal recipe. Keep exploratory testing in the strategy too.

Choose a layer by the failure boundary

Start with the risk, then ask what is the narrowest test scope that can detect it reliably. A test should exercise the boundary where the failure could occur, without adding layers that repeat the same assertion without adding confidence.

Layer Good fit Strength Cost or caution
Unit or component Local rules and behavior within an isolated unit or small component boundary Fast feedback, easier fault localization, and generally lower resource use Excessive simulation or narrow isolation can diverge from integrated behavior. Teams also use “unit” differently.
Integration, service, or API Contracts and interactions between components, services, databases, or dependencies Tests realistic interactions without requiring every check to traverse the full UI and system Requires more setup and resources than focused tests; boundaries vary by team.
End-to-end or UI Core user journeys where confidence depends on the assembled system and user-facing path Validates a high-fidelity flow through the system Can be slower, more expensive to maintain, and more exposed to timing, browser, and dependency instability.
Exploratory or manual Usability and unexpected quality concerns that are difficult to express as repeatable assertions Human-directed investigation can uncover gaps automation misses Does not provide the same repeatable regression check as an automated test.

These labels are not a universal taxonomy. Before comparing “unit” and “integration” counts, document what dependencies, processes, and interfaces each category actually exercises. Likewise, UI and end-to-end describe different characteristics: a UI test is not automatically an end-to-end test, and a test can validate an end-to-end path without defining its scope by the interface alone. Martin Fowler discusses both the value and the limits of the pyramid in The Practical Test Pyramid and Test Pyramid.

Balance speed, reliability, fidelity, and maintenance

Do not judge the strategy by test counts alone. Google’s SMURF model names five dimensions to consider: speed, maintainability, utilization, reliability, and fidelity. A focused test may run quickly and be easy to diagnose, while a broader test may offer stronger evidence about real interactions. Improving one dimension can affect another. Adam Bender’s Google-adapted overview, SMURF: Beyond the Test Pyramid, provides a framework for making those trade-offs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Scope and failure boundary: Does the test cross the boundary where the risk can occur?
  • Feedback speed: How soon will a developer learn whether a change broke the behavior?
  • Reliability: Does the test fail for useful reasons, or does instability obscure real regressions?
  • Fidelity: How closely do the exercised interfaces and dependencies resemble the operating conditions that matter?
  • Resource use and maintenance: What does the test require to run, debug, and keep current?
  • Added confidence: Does this higher-level check catch a risk the lower-level suite cannot?

The pyramid assumes broad tests tend to be slower, more brittle, and costlier than focused ones, but that is a heuristic, not a law. If your high-level tests are fast, reliable, and inexpensive to change, the right balance for your suite may differ.

Use 70/20/10 only as a starting guess

Google’s Mike Wacker suggested a first-guess split of 70% unit, 20% integration, and 10% end-to-end tests in a 2015 article, Just Say No to More End-to-End Tests. Google explicitly describes the ratio as a starting point whose exact mix differs by team. It is not a measured universal optimum, a required target, or a substitute for examining what your own failures and costs show.

Think of the pyramid as a prompt to keep broad checks selective, rather than a quota to fill. Disagreement about whether a suite should look like a pyramid, honeycomb, or another shape often reflects different definitions of the layers. Inspect the actual scope of tests before treating the shape as meaningful.

Build the portfolio around risk

  1. List important outcomes and failure modes. Identify the user outcomes that matter, component boundaries, external dependencies, and the risks that could harm those outcomes.
  2. Find the narrowest reliable detector for each risk. Put local rules in focused tests, boundary interactions in service or integration tests, and whole-system risks in end-to-end scenarios.
  3. Write down your team’s layer definitions. State which dependencies, processes, and interfaces each layer exercises so that category names do not hide different assumptions.
  4. Protect a small set of critical journeys. Add an end-to-end check when it covers a meaningful whole-system risk that lower layers do not cover. Avoid re-running every lower-level edge case through the UI.
  5. Use failures to improve the suite. When a broad test finds a defect, add a narrower regression check where practical, so later feedback is quicker and the defect is covered at the level closest to its cause.
  6. Retain exploratory testing. Investigate usability and unexpected behavior that is hard to express as automated assertions. Decide whether each discovery merits a repeatable test.

Fowler’s practical guidance emphasizes keeping end-to-end coverage purposeful, sequencing for useful feedback, and avoiding duplicated assertions. The UK Home Office’s Test pyramid guidance, last updated 31 October 2025, says its teams must avoid large numbers of end-to-end tests; that is a department-specific standard, not a universal regulation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Sequence tests for earlier, useful feedback

Run quick, narrowly scoped checks early when doing so helps developers find problems sooner. Run broader, slower checks later or at appropriate pipeline stages. Do not use a test’s name alone to decide when it runs: a fast, reliable integration test may belong early, while a costly check may need a different schedule. Choose sequencing based on the suite’s actual runtime and value.

Measure suite health over time using execution time, unreliable or flaky tests, defects discovered or escaping at each level, defect density, and automation coverage. The sources do not establish universal target values for these measures. Use them to identify local problems and trade-offs, not to chase a generic percentage.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

ScreenshotNeo for browser-level checks

For browser-level checks that need a captured page as evidence, ScreenshotNeo offers a screenshot API and MCP server. A screenshot can help inspect the rendered state of a page, but it does not replace assertions about application behavior, component boundaries, or user flows. Use it where a visual capture answers a defined test or debugging question.

Or skip the browser setup

A single GET request can return a screenshot or PDF. For example, this cURL request captures a page as WebP:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options and setup. ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can each be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for 1,000 free screenshots a month, with no card required.

Frequently Asked Questions

Should every important behavior have both a unit test and an end-to-end test?

No. Add a broader check when it covers a distinct integration or whole-system risk, rather than duplicating lower-level assertions without additional confidence.

Does a screenshot prove that a user journey works?

No. A screenshot captures rendered output at a point in time; it does not by itself establish that the underlying interactions or workflow behaved correctly.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.