iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A UI automation failure after a feature ships is a reason to investigate, not proof that the feature is broken. The cause may be a real regression, a selector invalidated by an intentional UI change, timing or environment instability, or a gap in the checks that ran before release. Reproduce the failure on the same build, classify what changed, and verify the user outcome before deciding whether to fix the product, the test, or the release process.
First determine whether the failure is repeatable
Run the failing test against the same build and environment, then compare the original run with a retry. Microsoft defines a flaky test as one that passes and fails nondeterministically on the same code in the same environment. Azure Pipelines also identifies the product, test code, environment, and flakiness as possible sources of failure. A retry that passes therefore does not establish that the original failure was harmless.
- Fails consistently: Treat it as a likely regression or a consistently invalid test assumption. Inspect the changed user flow and the assertion.
- Alternates between pass and fail: Investigate timing, shared state, test data, and environmental dependencies before attributing the result to a product defect.
- Fails only in a particular environment: Compare configuration, data, browser or device setup, and feature-flag state with a passing environment.
Use the first-run and retry results alongside logs, screenshots, and traces. Playwright’s retry guide explains how retries classify tests as flaky when a retry passes; that classification is a diagnostic signal, not a reason to ignore the initial failure. See Playwright’s test-retry guidance.
Check whether the UI change broke a selector or the user workflow
A UI test can fail even when users can still complete the intended task. Tests that locate elements through CSS structure or other implementation details are coupled to the DOM; a legitimate redesign or refactor can invalidate those selectors without changing the user-facing behavior. Conversely, a changed label, role, navigation path, or loading state can reveal a genuine usability or functional regression.
#1 Best Overall
- Review the feature diff for changed labels, roles, accessible names, DOM structure, test IDs, navigation, loading behavior, and feature-flag behavior.
- Identify the failing locator and assertion. Ask whether the test is checking what a user can see and do, or a particular implementation detail.
- Verify the intended user outcome manually or through a focused test using the current UI contract.
- If the workflow still works, update a brittle selector to match the intended contract. If the outcome is missing or wrong, investigate the product change.
Playwright recommends user-facing locators and explicit contracts, noting: “To make tests resilient, we recommend prioritizing user-facing attributes and explicit contracts.” Its locator guidance explains that DOM changes can invalidate selectors tied to page structure. Prefer accessible roles, labels, and meaningful text when they represent the actual interface; use an explicit test ID when the team needs a stable, intentional test contract that is not well represented by user-facing attributes. See Playwright’s best practices.
Separate asynchronous timing from a missing outcome
Interfaces often update asynchronously: a button may become enabled after data loads, a confirmation may appear after a request, or navigation may complete after an action. A test that checks too early can fail despite a valid workflow. Framework waiting helps with that timing mismatch, but it cannot make a wrong assertion correct or make an outcome appear when the product never produces it.
Playwright’s documentation says, “Locators come with auto waiting and retry-ability.” Its actionability checks and retrying assertions wait for the relevant conditions; if the condition never becomes true within the timeout, the test still fails. Use those mechanisms to wait for meaningful outcomes, such as a visible confirmation or expected destination, rather than adding arbitrary pauses. Keep timeouts useful and diagnostics available so a genuinely absent outcome remains visible. See Playwright’s actionability guidance.
Trace the failure through the release path
A test can only prevent a bad release if it runs at the right point and its result affects promotion. Check which tests ran for the shipped build, whether failures gated deployment, and what health signals were checked afterward. Microsoft Azure Well-Architected guidance warns that “Failed deployments and erroneous releases are common causes for application outages.” It recommends continuous validation and testing as part of release safety. See Microsoft’s continuous integration and delivery guidance.
Rank #3
- Before promotion: Confirm that relevant UI tests ran against the build being released and that their results were visible to the people or automation approving deployment.
- During rollout: Where practical, use a feature flag or gradual exposure to limit the number of users affected while validating the change.
- After deployment: Monitor workload health and user-facing behavior, and know what signal triggers rollback. AWS Well-Architected guidance treats post-deployment health monitoring as part of safe deployment, alongside testing, rollback, and feature flags. See AWS’s deployment-risk guidance.
Passing tests before release and healthy production behavior answer different questions. Pre-deployment checks assess known scenarios; post-deployment monitoring can reveal issues those scenarios did not cover.
Keep flaky tests visible while repairing them
If a test is intermittent, record its classification, owner, evidence, and follow-up rather than silently excluding it. A quarantine can keep an unstable test from blocking unrelated work while the cause is addressed, but it is useful only when failures remain visible and the test has a path back to blocking status after repair. Microsoft’s account of its own test-management system describes a quarantine process that tracks tests and restores them after bugs are resolved. It also reported identifying approximately 49,000 flaky tests and helping 160,000 sessions pass; those are Microsoft-specific figures from 2022, not general industry rates. See Microsoft Engineering’s account of flaky-test management.
Rank #4
Internal case reports can illustrate what a team has achieved, but they are not universal benchmarks. Delivery Hero Tech reported more than 900 UI test cases and an approximately 98% monthly average success rate through December 31, 2024, in its account of internal mobile testing. The same account described about 60% of testing builds as flaky. These are organization-specific retrospective figures, not estimates of typical UI test performance. See Delivery Hero Tech’s testing account.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Turn the incident into a concrete follow-up
Once the failure is classified, assign the right fix rather than applying a blanket retry or timeout increase.
Quick Recap
- Product regression: Repair the broken user outcome, add or correct an assertion for it, and verify the relevant release gate.
- Selector broken by an intentional UI change: Update the test to target the new user-facing contract or an explicit test contract, then consider whether the change should have been covered by an affected test.
- Timing or nondeterminism: Remove shared-state or data dependencies where possible; wait for a meaningful condition; capture diagnostics; and track the test until it is reliable.
- Coverage or release-path gap: Add the missing scenario at an appropriate test level, ensure it runs against the release candidate, and connect results to promotion and post-deployment health checks.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

