Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
There is no universal number of weeks for an A/B test. Estimate how many eligible users each variant needs to detect the smallest effect that would change your decision, convert that sample into elapsed time using realistic traffic, and follow a stopping rule suited to your analysis method.
What determines how long your test needs to run?
Duration is the result of a sample-size requirement and the rate at which eligible users enter the experiment. The sample requirement depends on your primary metric, its baseline behavior, the effect you want to detect, and your chosen statistical decision criteria. More eligible traffic can shorten the calendar time; a smaller effect to detect generally increases the required sample.
Use the population that can actually enter the test, not total site traffic by default. A test limited to a particular region, device, audience, or page type may receive only a fraction of overall traffic. Statsig’s power-analysis documentation discusses planning around the experiment population, while Amplitude’s key terms explain sample size and related experimentation concepts.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Plan the test before estimating its duration
- Write the hypothesis and choose one primary metric. Identify separate guardrail metrics for outcomes that must not worsen, such as errors or cancellations. A clear primary metric makes it less tempting to keep checking many outcomes until one looks favorable.
- Set the minimum detectable effect (MDE). This is the smallest change worth acting on—not simply the smallest change your analytics might detect. Amplitude’s MDE guidance explains how the chosen effect size affects planning. If you want to detect a smaller effect while holding other assumptions constant, you will generally need more observations.
- Choose the analysis and decision criteria. For a fixed-horizon test, specify the target sample and statistical criteria before the test starts. For a sequential test, define the interim decision procedure and its criteria in advance.
- Estimate sample needs with inputs for the eligible population. Use a relevant baseline rate or metric mean and variance, along with the planned allocation between variants. Statsig describes these inputs in its power-analysis documentation. If only a subset of users is eligible, use that subset’s expected exposure rather than total traffic.
- Translate the sample into calendar time. Estimate how many eligible exposures each variant receives per day under the planned allocation, then calculate how long it will take to reach the target. Amplitude’s duration-estimation documentation describes estimating run time from inputs such as means, variances, and exposure rates.
- Check whether the calendar adds a constraint. Consider weekly patterns, delayed conversions, seasonal drift, or a product system that needs time to learn. Add enough calendar coverage for those conditions to be represented; do not assume reaching the sample target alone resolves them.
Choose a stopping rule that matches your method
Fixed-horizon testing
A fixed-horizon plan commits to a sample target and evaluates the result at the planned endpoint. Repeatedly checking ordinary fixed-horizon significance and stopping as soon as a result looks favorable can inflate false-positive risk. If you choose this method, do not treat interim significance as permission to end early; use the planned endpoint and analysis.
#1 Best Overall
Sequential testing
Sequential methods adjust the analysis to support interim decisions. They can be appropriate when you need to review evidence during the run, but they are not a blanket license to stop whenever a dashboard changes color. The sequential procedure must be active and applied as intended, and an early result on the primary metric does not establish that guardrails had enough data to rule out harm. See Statsig’s frequentist sequential-testing documentation for its method and setup.
Why a duration estimate can change
A calculated run time is a forecast, not a guarantee. It relies on assumptions about exposure rates and metric behavior. Amplitude notes that duration estimates use inputs including means, variances, and exposure rates; seasonality or changes in those inputs can reduce accuracy. One estimator workflow assumes constant daily exposure. Revisit the estimate if traffic, targeting, allocation, variance, or the metric’s behavior changes materially.
Longer is not automatically better. Once the planned evidence and practical decision criteria are met, continuing without a reason can expose the experiment to changing conditions and delay a decision. For website experiments, Google Search Central advises ending testing after enough data has been collected and removing test elements after the experiment concludes. Its guidance is specifically about website testing and search; it is not a universal sample-size formula. Read Google Search Central’s A/B testing best practices.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhen does the four-week recommendation apply?
Google Ads API documentation recommends running its campaign experiments for at least four weeks to account for weekly cycles, conversion delays, and learning periods. That is product-specific guidance for Google Ads campaign experiments—not a general rule for website, app, or product A/B tests. Use it when planning within that context, not as a substitute for calculating sample needs elsewhere. The guidance appears in Google Ads API reporting documentation.
Quick Recap
Rank #4
Rank #3
How to decide that the test is finished
- The planned sample or valid sequential decision criterion has been reached.
- The result is interpreted against the preselected primary metric and practical MDE, rather than chosen after seeing the data.
- Guardrails have sufficient evidence for the decision you need to make; early significance on one metric alone may not establish this.
- Relevant calendar patterns, conversion delays, or learning periods have been accounted for.
- For a website test, remove the test markup or code after the test ends, as Google Search Central advises.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

