Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A Playwright test that fails and then passes on retry is classified as flaky: the runner has identified an inconsistent result, not confirmed that the underlying problem is fixed. To find the cause, inspect the first failure, check what the test waits for, verify locator behavior and test isolation, then use a CI trace to see what happened. Retries can expose intermittence and help capture evidence, but they are not a repair by themselves.

What a “flaky” result actually tells you

Playwright’s test runner reports a test as flaky when it fails on an attempt and passes on a retry. That label describes the outcomes; it does not identify whether the cause was timing, shared state, an unreliable selector, or something else. A green retry therefore does not make the first failure irrelevant. Start with the failed attempt, not just the final green status. Playwright documents retry behavior and flaky classification.

Diagnose the failure in this order

1. Separate the first attempt from the retry

Use the test report to determine whether the test passed immediately, failed consistently, or failed first and passed on retry. The last pattern is evidence of intermittence. Record the failed assertion or action and its error before changing anything; otherwise, a later pass can obscure the behavior you need to explain.

2. Check whether the test waits for the condition it cares about

Browser interfaces update asynchronously. A plain getter or immediate boolean check can sample the page before the expected state appears. For expected UI states, use a web-first assertion such as await expect(locator).toBeVisible(). Playwright rechecks the condition until it succeeds or times out; the documented default assertion timeout is five seconds, and it can be configured in test settings or for an individual assertion. See Playwright’s assertion guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Actions and assertions wait at different points. A locator action such as click() waits for the target to be actionable before interacting; an assertion waits for the resulting UI state you expect. Neither substitutes for the other. For a complex condition, Playwright also documents expect.poll and expect.toPass. Configure toPass deliberately: its default timeout is zero, and it does not use the custom expect timeout. The assertion documentation covers these options.

3. Inspect the locator and the action’s timeout

Before a click, Playwright checks that the locator resolves to one element and that the element is visible, stable, able to receive events, and enabled. If the action times out, that is useful evidence: one or more of those conditions may never have become true. Read the error and inspect the target instead of assuming that adding time will solve it. Playwright lists the actionability checks.

Prefer selectors that describe the user-facing interface: a role and accessible name, a label, or meaningful text. Use a test ID when you need a deliberate testing contract rather than a suitable user-facing locator. Avoid selectors tied to incidental implementation details when a more resilient option is available. Playwright’s locator guide describes these choices.

Do not make force: true a routine flakiness fix. Force disables non-essential actionability checks, including the check that the target receives events. That can hide an overlay or other interaction problem instead of resolving it. The actionability guide explains force behavior.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Run the test alone and look for hidden shared state

Playwright recommends independent tests. Each test gets an isolated browser context with its own browser state, including cookies and storage, but isolation does not make a test independent if its setup relies on data or effects created by another test. Assumptions about test order, shared external data, or setup that is not repeated for each test can make a failure hard to reproduce. Playwright’s best practices recommend isolation; its browser-context guide explains isolated state.

  • Run the failing test by itself and see whether the failure remains.
  • Check whether it assumes another test has created or changed data.
  • Check whether cookies, local storage, or other setup are being relied on rather than established by the test.
  • Make required setup self-contained so the test can be rerun without depending on order.

5. Capture a trace from CI and inspect the failed attempt

A trace provides a timeline of actions and DOM snapshots that you can inspect in Playwright Trace Viewer. It can help distinguish a slow interface from an intercepted click, an unexpected page state, or an assumption that was never true. Playwright recommends capturing traces on the first CI retry; recording a trace for every test can add substantial performance cost. See the Trace Viewer guide.

Trace configuration is selectable: documented options include on-first-retry, on-all-retries, and retain-on-failure. The appropriate setting depends on the diagnosis you need and the cost of collecting artifacts; check the configuration documentation for the options available to your Playwright version. Playwright’s test configuration reference covers retries and tracing.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose a fix that addresses the evidence

Once you have a reproducible failure or trace evidence, choose the change that matches it. A longer timeout may be appropriate if the intended condition genuinely takes longer, but it can also delay detection without fixing a race or incorrect expectation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Evidence or suspected cause Targeted response What to verify
A manual check samples the page before the expected state appears Use a web-first assertion for the expected UI condition The assertion waits for the right state, and the failure message remains useful
A click times out because the target is hidden, moving, disabled, or does not receive events Address the actual UI or locator condition; do not routinely bypass checks with force The target becomes actionable through the normal user interaction
The test passes alone but depends on another test’s data or order Make setup self-contained and remove the order dependency The test can run alone and alongside the rest of the suite
The trace shows a state different from the one the test assumes Correct the setup, expectation, or application behavior responsible for that state The original failure condition no longer occurs, rather than merely passing on retry
No cause is established yet, but CI intermittence needs investigation Keep retries intentional and collect a trace on retry The first failed attempt remains available for inspection

Why “it passes locally” is not a diagnosis

A local pass shows that the test succeeded under that run’s conditions; it does not establish why a CI attempt failed. Compare the failed CI action, expectation, state, and trace with the local run rather than presuming one universal environmental cause. The documented Playwright mechanisms explain how waiting, isolation, retries, and tracing work, but they cannot identify the cause in a particular project without its failure evidence.

Use retries as a signal, not a substitute for reliability

Retries are useful when you need to reveal intermittent outcomes and capture diagnostic artifacts. Keep them configured intentionally, and treat a failure followed by a pass as a prompt to investigate. After changing the test or application, verify the original failure condition and confirm the test still runs independently; do not treat a larger retry count or a green rerun alone as proof of a fix.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.