Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Headless testing is reliable when it tests user-visible behavior in isolated browser contexts, runs an intentional browser matrix, and uses deterministic CI settings. Treat headless mode as a way to run a real browser without displaying its window—not as a different kind of application. The same cookies, JavaScript, network failures, responsive layouts and browser differences still affect your users.

The practices below show how to build a maintainable Playwright suite, where Selenium still fits, how to control parallel CI execution, and how to diagnose failures without turning every test into a slow recording session.

What headless testing is—and what it is not

In headless mode, Chromium, Firefox or WebKit executes without a visible desktop window. It is useful on CI workers and containers, but it does not remove the need for realistic viewport, device and browser coverage. A passing Chromium run cannot prove that a WebKit user, a mobile-sized viewport or a user with a different storage state will see the same result.

Use headless end-to-end tests for functional questions such as: can a visitor sign in, submit a form, complete checkout, navigate with the keyboard and see the correct error? Do not use them as a substitute for load testing. Selenium’s documentation says performance testing with Selenium/WebDriver is generally not advised because browser startup, servers, third-party resources and WebDriver instrumentation introduce uncontrolled variation. Measure throughput and resource behavior with a dedicated performance tool, then use browser tests for a small set of user-critical checks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Eight practices that prevent flaky headless tests

1. Assert what a user can see and do

Playwright recommends verifying end-user behavior rather than implementation details such as function names, array structure or CSS classes. Prefer accessible roles, labels and visible text; use a test id only when a stable user-facing locator is not available.

import { test, expect } from '@playwright/test';

test('customer can submit a contact form', async ({ page }) => {
  await page.goto('/contact');
  await page.getByRole('textbox', { name: 'Email' }).fill('sam@example.com');
  await page.getByRole('textbox', { name: 'Message' }).fill('Please call me.');
  await page.getByRole('button', { name: 'Send message' }).click();
  await expect(page.getByRole('status')).toHaveText('Message sent');
});

This style survives refactoring better than selectors tied to a component’s internal class names. It also exposes accessibility problems early because a missing label can make a locator fail for the same reason it would be difficult for a user to operate.

2. Isolate every test

Each test should receive independent cookies, local storage, session state and data. Playwright’s isolation model uses separate browser contexts and states that isolation improves reproducibility, debugging and protection against cascading failures.

  • Create a user or account fixture for the test instead of reusing a mutable shared account.
  • Reset or uniquely namespace database records before the test starts.
  • Do not depend on test order, a previous test’s local storage, or a server-side record left behind by another worker.
  • Use a dedicated setup project for immutable authentication state, then give tests their own context and data.

3. Choose browser projects deliberately

Build a matrix from your audience and risk, not from a checkbox. Playwright documents cross-browser projects for Chromium, Firefox and WebKit; its browser guidance is at playwright.dev/docs/browsers.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Project What it represents When it is essential
Chromium Chrome and Chromium-based desktop users Most teams’ baseline and a useful fast feedback project
Firefox Firefox desktop users Sites with standards-sensitive layout, extensions or Firefox traffic
WebKit Safari’s browser engine behavior Products with iPhone, iPad or macOS Safari users
Branded Chrome or Edge Enterprise or policy-managed installations When your support contract names a branded browser
Device profile Mobile viewport, touch and device characteristics Responsive flows, mobile navigation and checkout

Run the smallest representative set on every pull request and the full matrix on a schedule or before release. Add a project only when it represents a real user segment or a known compatibility risk.

4. Make CI deterministic

Set an explicit global timeout so a hung test stops cleanly. Choose workers based on the CPU and memory available to the job; more workers are not automatically faster when the application, database or runner is contended. Install only the browser binaries required by that job. Linux is often the economical CI choice, but validate rendering and fonts on the operating systems you support.

import { defineConfig, devices } from '@playwright/test';

export default defineConfig({
  testDir: './tests',
  timeout: 30_000,
  expect: { timeout: 5_000 },
  workers: process.env.CI ? 2 : undefined,
  retries: process.env.CI ? 1 : 0,
  reporter: [['list'], ['html', { open: 'never' }]],
  use: {
    baseURL: 'http://127.0.0.1:3000',
    headless: true,
    trace: 'on-first-retry',
    screenshot: 'only-on-failure',
    video: 'retain-on-failure'
  },
  projects: [
    { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
    { name: 'firefox', use: { ...devices['Desktop Firefox'] } },
    { name: 'webkit', use: { ...devices['Desktop Safari'] } }
  ]
});

The timeout, worker count and retry policy are explicit so a change in the runner does not silently change test behavior. Keep the browser installation step pinned to the Playwright version used by the project.

5. Parallelize only after data is safe

Playwright runs test files in parallel by default, with separate worker processes and isolated browser contexts. Parallel execution is valuable only after tests can run in any order. If workers compete for a single account, a shared cart, a rate-limited API or a small database, reduce the worker count first. Once the suite is independent, shard it across machines to shorten wall-clock time:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
npx playwright test --shard=1/4
npx playwright test --shard=2/4
npx playwright test --shard=3/4
npx playwright test --shard=4/4

Collect the report and artifacts from every shard. A green shard is not a green build if another shard failed.

6. Wait for conditions, not arbitrary sleeps

Use locator assertions and Playwright actions that wait for the element to be actionable. For asynchronous screens, wait for a specific user-visible condition such as a heading, status message or table row. A fixed sleep hides slow environments and wastes time on fast ones. When a test truly depends on a background transition, expose a deterministic application signal or wait for a narrowly scoped network response rather than adding a large delay.

7. Trace on failure or retry

Playwright recommends collecting a trace on the first CI retry rather than every test because always-on tracing is performance-heavy. A trace contains a timeline, DOM snapshots and network information, allowing you to see what the page looked like immediately before the failure. Preserve the HTML report, trace, screenshot and video for failed runs, and publish them as CI artifacts with a retention period that matches your debugging needs.

8. Keep the test toolchain current

Update the Playwright package and its browser binaries together. Review release notes before updating, run the matrix, and commit any required snapshot or baseline changes deliberately. TypeScript and ESLint help catch mistakes; enable @typescript-eslint/no-floating-promises so a missing await does not turn an action into a race.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A minimal CI workflow

The following GitHub Actions job installs dependencies, installs only Chromium, starts the application and preserves the report when a test fails. Adapt the start command and package-manager cache to your project.

name: browser-tests
on: [push, pull_request]
jobs:
  e2e:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with:
          node-version: 20
          cache: npm
      - run: npm ci
      - run: npx playwright install --with-deps chromium
      - run: npm run build
      - run: npm run start -- --host 127.0.0.1 &
      - run: npx playwright test --project=chromium
      - if: always()
        uses: actions/upload-artifact@v4
        with:
          name: playwright-report
          path: |
            playwright-report/
            test-results/

Use a separate job or scheduled workflow for Firefox, WebKit and branded-browser coverage if the pull-request budget is limited. The matrix should still run before a release.

Playwright or Selenium?

Neither tool is universally best. Selenium’s own test-practice guidance notes that no single approach works for every situation. Decide using the workload and the infrastructure you already operate.

Decision axis Playwright Selenium WebDriver
Browser-engine coverage Chromium, Firefox and WebKit projects are configured in one runner. Broad browser and vendor ecosystem through WebDriver implementations.
Isolation Separate BrowserContexts are a first-class model for test isolation. Isolation is assembled with driver sessions, profiles and your fixtures.
Waiting and diagnostics Locator waiting, trace viewer, DOM snapshots and network details are integrated. Waiting, logging and diagnostics depend more on your language bindings and framework.
CI scaling Worker controls and sharding are built into the test runner. Scaling commonly follows the grid, runner and orchestration system you already use.
Language and legacy fit Strong fit for teams adopting the Playwright runner and its supported languages. Often the practical choice when an organization already has WebDriver grids, bindings or shared fixtures.
Performance measurement Use for functional user journeys, not load generation. Selenium documentation also advises against using WebDriver as a performance-testing tool.

Whichever runner you select, keep assertions user-facing, isolate state and make CI resources explicit. Changing frameworks will not repair tests that share data or rely on timing accidents.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical failure-investigation workflow

  1. Classify the symptom. Decide whether the failure is a product defect, an environment problem, a locator problem, a data collision or a timeout.
  2. Open the first retry trace. Inspect the action timeline, DOM snapshot and network requests at the failed step.
  3. Reproduce with one worker. Run the failing project and test alone. If it passes only when isolated, investigate shared data or resource contention.
  4. Check the browser matrix. A failure limited to WebKit or Firefox may indicate a genuine engine difference rather than flakiness.
  5. Remove timing guesses. Replace sleeps with a visible assertion, an actionability check or a narrowly scoped response wait.
  6. Fix the fixture. Create unique records and clean them up; do not add retries to conceal a polluted environment.
  7. Retain evidence. Upload the trace, screenshot, video and report from the failed run so another engineer can diagnose it without rerunning immediately.

Reliability, speed and cost trade-offs

  • Fast feedback: run Chromium smoke tests on pull requests, with a modest worker count that the CI machine can sustain.
  • Coverage: run Firefox, WebKit, device profiles and branded browsers according to actual support commitments.
  • Resource control: parallel workers consume CPU, memory, database connections and service quota. Reduce workers when failures correlate with runner pressure.
  • Artifact cost: traces and videos are valuable on failure but expensive to collect for every test; the first-retry policy keeps evidence targeted.
  • Maintenance: browser and dependency updates are part of test ownership. A stale binary can report failures that users no longer see, while an unreviewed update can change rendering or timing.

Or skip the browser setup

If your requirement is a clean website image or PDF rather than an assertion-driven test, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

One request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options and response details. The same call from Python is:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. It supports full-page and element captures, lazy-image loading, dark mode, device and viewport settings, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently asked questions

Frequently Asked Questions

Does headless mode test visual appearance?

It renders the page, but functional assertions alone do not establish that pixels match a design. Add a deliberate visual-regression process with reviewed baselines if visual fidelity is a requirement.

Should every test run in every browser?

No. Map projects to supported user segments and risk, then run a small smoke set broadly and deeper journeys where the business impact justifies the cost.

When is a retry acceptable?

A single CI retry is useful for collecting a trace and distinguishing transient infrastructure failures. It should not replace fixing shared data, unstable locators or arbitrary waits.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.