iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Cutting regression testing from weeks to days is usually a sequence of changes, not a single switch. Start by making the existing suite faster to execute: run independent tests in parallel, fix the infrastructure that starves them, and remove stale or ineffective tests. Then run only the tests tied to a change, order the rest so likely failures surface first, and keep a slower full-suite run on a schedule. Published “weeks to hours” stories exist, but they describe specific systems, so the realistic target is one you can measure on your own pipeline.
Before choosing any technique, confirm what “weeks” means. If it is elapsed wall-clock time for a full run, queueing, serialized environments, and slow setup often dominate, and selection will not fix them. If it is the time a developer waits for a first useful signal, ordering and selection matter more than total runtime. The two problems call for different fixes.
The main approaches and what each one trades
Several real techniques address regression time, and they differ in what they change, how much they can delay or omit, and how much evidence they need. Compare them on wall-clock improvement, the scope of tests delayed or skipped, fault-detection evidence, implementation and maintenance cost, dependence on historical data, test-dependency safety, and how critical the system is.
| Approach | What it changes | Useful when | Main caution |
|---|---|---|---|
| Parallel execution | Runs independent tests concurrently | Total wall time is high and workers or environments can scale | Shared state and dependencies can make parallel runs unreliable; watch resource contention. See Microsoft Learn guidance on testing practices. |
| Suite cleanup | Removes stale or duplicate tests and repairs ineffective or flaky ones | The suite has accumulated test debt or low-signal checks | Do not delete a test just because it is slow; first confirm the behavior and risk it covers. |
| Test ordering | Runs likely failures earlier | The full suite must still run, but feedback should arrive sooner | Ordering alone does not necessarily reduce total completion time. Measure time to first failure. |
| Change-based selection (test impact analysis) | Chooses tests related to modified code | Code-to-test relationships can be derived and maintained | Missed dependencies can drop relevant checks, so keep broader runs elsewhere in the pipeline. |
| Predictive selection | Uses historical changes and results to predict relevant tests | Reliable historical data exists and the risk can be governed | Model uncertainty is real. AWS advises against excluding security tests or relying on predictive selection for sensitive critical systems. See AWS DevOps Guidance on advanced test selection. |
| Time-budgeted prioritization | Stops a prioritized run at a chosen time limit | You can quantify failure yield and accept an explicit risk | A locally chosen budget may miss failures. Keep full-suite coverage elsewhere and re-evaluate the budget as the suite changes. |
A staged sequence for cutting run time
AWS guidance gives a sensible order. It advises that, before implementing advanced test selection with machine learning, teams should first optimize test execution through parallelization, reducing stale or ineffective tests, improving the infrastructure the tests run on, and changing the order of tests to optimize for faster feedback. The steps below follow that logic.
1. Establish a baseline
Record the numbers that tell you where time goes before changing anything:
- Wall-clock duration of the full run and of the pipeline stage that contains it
- Queue time, separated from execution time
- Time to first useful failure
- Total test count, failure rate, and flakiness rate
- The areas of the product each test covers
Separate slow tests caused by the test itself from those waiting on shared infrastructure or serialized resources. A test that sleeps for a fixed interval needs a different fix than one stuck behind a single database instance. Azure guidance recommends monitoring execution-time trends and test reliability measures over time, and Shopify uses time to first failure and fault-detection measures to judge prioritization. A baseline lets you tell which change actually moved the numbers.
2. Remove avoidable execution waste
Parallel execution is the first lever when tests are independent. Scale workers or improve the test environment if worker scarcity or setup delays dominate. Review stale, obsolete, and duplicate tests, and repair unreliable ones rather than assuming they provide meaningful assurance.
Parallelism shortens elapsed time without reducing the number of tests, but it does not remove shared state or hidden dependencies. Partition and order tests with those dependencies in mind. A 2020 ISSTA abstract on dependent-test-aware regression testing techniques warns that dependence can contribute to flaky failures when tests are reordered, selected, or parallelized, so a faster run that turns green-to-red randomly has not really saved time.
3. Select tests related to the change
Change-based test impact analysis examines the code difference and identifies tests likely to be affected. AWS describes this as a structured way to run a relevant subset without machine learning. Google’s 2014 work on regression testing in continuous integration describes selecting tests before a change is submitted and testing dependent modules after submission, which splits the work across two points in the workflow.
Maintain the selection map as the architecture and test coverage evolve. An unselected test is delayed, not proven irrelevant forever. That distinction should shape how you schedule the rest of the suite.
4. Prioritize the selected tests
Prioritization changes the order of tests, while selection changes which tests are in the run. Put tests with a stronger historical failure signal, or direct relevance to the change, at the front, so a failing change is discovered sooner. Shopify’s approach layered a history-based prioritized set on top of change-based selection and measured performance under fixed time limits.
5. Add a time budget only after measurement
A time budget stops a prioritized run at a chosen limit. Set that limit from the observed distribution of failure yield and from the risk your team accepts, not from an arbitrary target. Shopify’s 2022 engineering write-up on time-constrained CI feedback is a useful model of how to quantify that trade-off. Its figures describe its own system, and they are discussed below.
6. Keep a slower full-suite safety net
Use the fast checks for frequent feedback, and reserve slower integration, load, performance, or broad regression suites for nightly, pre-release, or other appropriate stages. Azure guidance recommends nightly full-suite runs in pre-production for long-running tests and fail-fast handling for critical tests. AWS recommends running a full set asynchronously when predictive selection is used, so that eventual full results still arrive for every change.
Rank #4
7. Review results as quality signals
Track execution-time trends alongside pass rate, coverage, flakiness, and defect escape rate. If a production defect escapes, add or correct a regression test where the gap occurred. Avoid making coverage percentage the only target. Azure guidance treats it as a signal and asks teams to emphasize high-risk paths.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What published results show, and what they do not
Published numbers are useful for setting expectations, but each one is tied to a system, a measurement, and a date. Read them with those qualifications attached.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →- Shopify, 2022 (test budget analysis). In the mean case, the failure-rate prioritization criterion found 80% of failures after running 60% of the selected tests. In its 5th-percentile view, 70% of the selected suite found 50% of failures. The selected suite was a median 40% of the full suite. These figures describe Shopify’s large monolith and data, and should be re-measured locally before use.
- Microsoft Research, 2020 (flaky tests). In an evaluation of five tests affected by asynchronous calls, the proposed FaTB approach reduced runtimes by up to 78%. The paper reports no empirical change to how often those tests failed flakily in that evaluation, and five tests is a small sample. Do not assume the same gain across a suite.
- Di Nardo and colleagues, 2015 (industrial system). Test-suite minimization using finer-grained coverage reported 79.5% execution-cost savings with fault-detection capability above 70%. Test selection in the same study saved less than 2%. The contrast shows that methods perform very differently depending on the changes and the context, and minimization is not the same as selection.
- Perfecto case study, date not displayed (via CaseStudies.com). A vendor-attributed account describes an unnamed large North American bank that cut a 2,000-test suite from two weeks to seven hours using code optimization and parallel execution, with automated coverage stated as roughly 70% per release. The bank is unnamed, the account is vendor-attributed, and the result is a case study, not an independent benchmark or a typical outcome.
The most consistent pattern across these sources is that the largest gains came from combining execution improvements with selection or prioritization, and that reducing the run was never free of a coverage decision.
Best Value
Flaky tests inflate run time and mislead the team
Flaky tests cost time twice: they consume runner capacity, and they trigger reruns that hide real failures. Microsoft Research’s 2020 study, by Wing Lam, Kivanc Muslu, Hitesh Sajnani, and Suresh Thummalapenta, defines the problem this way: flaky tests, which nondeterministically pass or fail on the same code, are problematic because they provide misleading signals during regression testing. In the six studied Microsoft projects, asynchronous calls were a leading cause.
Treat flakiness as a reliability problem to diagnose, not as noise to rerun. Quarantine a flaky test only with a named owner and a deadline, and fix the root cause before the test returns to the blocking path. The FaTB result above is one example of a targeted approach, scoped to the tests it was evaluated on.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

