To run a useful email A/B test, define one question, change one variable, randomly divide eligible recipients between versions, and choose the main success metric before sending. Set the audience share and test duration, then follow the platform’s result rules. If the outcome is inconclusive, treat it as uncertain—not as a winner.
Plan the test before building the email
Write down the decision the test should inform: what you will change, which audience you will test, what outcome matters, and what you will do with the result. For example: “For this audience, does a benefit-focused subject line increase qualified clicks compared with our current subject line?” This is a hypothesis to test, not a predicted result.
Choose one variable at a time. Common email variables include the subject line, sender name, content, and send time. If you are testing a content element, keep the other email content the same; changing several things at once makes it difficult to tell which change drove the result. Klaviyo gives the same guidance in its A/B testing instructions, while Mailchimp outlines its available test types in About A/B Tests.
Choose the metric that answers your question
Pick a primary metric before the campaign goes out. Use the outcome closest to the campaign goal rather than selecting whichever metric makes one version appear stronger after the fact.
#1 Best Overall
- Subject line, preview text, or sender identity: Klaviyo commonly pairs these tests with open rate. However, opens can be inflated by Apple Mail Privacy Protection, so clicks or downstream conversions may be more useful if the real goal is engagement or sales.
- Email body, button, layout, or offer: Click rate is a closer measure of interaction with the message. For a purchase-focused question, use a conversion or placed-order metric if your account supports it and the measurement window is long enough.
- Send time: Keep content and subject lines equivalent so timing is the variable being tested. Use the result to inform future scheduling.
Klaviyo describes metric choices and availability in its campaign A/B test guidance. Placed-order measurement is not available for every variation type or account setup, so check the options in your platform before relying on it.
Set up the audience and versions
- Define the eligible audience. Use the same audience rules for both versions, and exclude recipients who should not receive the campaign.
- Create clearly named variations. Change only the chosen element. Keep the offer, audience criteria, and other content consistent; for a send-time test, timing is the intended exception.
- Use the platform’s audience assignment. Mailchimp describes randomly selecting contacts for test variations. Both Mailchimp and Klaviyo support dividing an audience among versions; consult the relevant setup flow for current controls and eligibility.
- Choose test share and duration. If your platform offers a staged test, decide what portion of recipients should receive each variation and when the platform should select a winner. Klaviyo allows test size and duration to be adjusted; when fewer than all recipients are in the initial test, the remaining audience can receive the selected winner.
Plan the send so that versions are comparable. Avoid changing audience criteria or campaign conditions between variations unless one of those conditions is the variable under test. Platform features and access can depend on account or plan; verify current availability in your account.
Rank #2
Send the test and interpret the result
After the test runs, review the primary metric and the platform’s uncertainty label, not just which version has the larger displayed number. Klaviyo’s reporting distinguishes significant, promising, not significant, and inconclusive outcomes. Its campaign-specific rule tags a result as statistically significant when each variation has at least 50 recipients and the win probability is at least 90%. These are Klaviyo product criteria, not universal sample-size or statistical-power rules. See Klaviyo’s explanation of statistical significance.
A larger list is not automatically enough to detect a small effect. For a custom test, the required audience depends on the baseline rate, the smallest lift worth acting on, the number of variations, and the error tolerance you want. The cited platform guidance does not provide one sample-size figure that works for every campaign. If the evidence is inconclusive, record that outcome and consider a better-powered retest rather than declaring a winner.
Rank #3
Account for privacy-related open inflation
Apple Mail Privacy Protection can prefetch tracking pixels, which may inflate recorded opens. Klaviyo notes that this can make open-rate significance harder to reach and recommends considering a custom report with an MPP property when a large share of opens comes from Apple Mail. This limitation affects open measurement; when clicks or conversions more directly answer your question, use those outcomes instead. Details are in Klaviyo’s A/B test guidance.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Apply the learning to this and later campaigns
If your test covers only part of the audience and the platform supports winner deployment, send the selected version to eligible recipients who have not yet received the campaign. Klaviyo describes this staged approach; Mailchimp’s subject-line workflow likewise describes sending the better-open-rate subject line to the remaining subscribed audience. For a test already sent to the full audience, the result cannot change that past send—use it to guide a later campaign instead. Mailchimp’s workflow and subject-line advice are documented in Best Practices for Email Subject Lines.
Keep a brief record of the hypothesis, audience, variable, metric, test share, duration, and result label. That makes it easier to distinguish a repeatable learning from an uncertain outcome and prevents a later campaign from unintentionally changing several things at once.
Quick Recap
Best Value
Quick pre-send checklist
- Is the hypothesis tied to a specific campaign decision?
- Does each variation differ in only one variable?
- Are audience rules and other conditions comparable?
- Is the primary metric chosen before results arrive?
- Are the test share, duration, and remaining-audience plan clear?
- Does the account support the metric and staged-send feature you intend to use?
- Will you treat an uncertain result as inconclusive rather than force a winner?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

