1,038 events / 12,500 units
95% interval 7.83%–8.80%Email A/B test
calculator
Compare outcomes without turning one threshold into a certainty machine. See the estimated effect, its plausible range, practical value, and data-quality checks.
- Effect intervals
- SRM diagnostic
- Sample planning
Design assumptions +
Use unique randomized units—not sends, opens, or repeated events—unless each observation is genuinely independent by design.
Evidence favors the variant; value remains uncertain
The 95% effect interval is above zero, but part of it remains below the minimum worthwhile effect.
5,120 planned units remain. Do not use an ordinary fixed-horizon threshold as an early-stopping rule.
1,127 events / 12,380 units
95% interval 8.61%–9.62%+0.80% absolute improvement
The interval crosses the 8.0% relative threshold. A worthwhile effect remains possible, but is not established by this sample.
Allocation looks plausible
Observed B share 49.76% versus 50.0% expected. SRM p-value 0.4468.
Large-sample check passes
Both arms have at least five observed events and non-events.
What the number does—and does not—support.
Repeatedly checking and stopping when a conventional p-value crosses a threshold inflates false-positive risk without a sequential design.
The two-sided p-value is 0.0253 under a no-difference model. It is not the probability that either variant is better.
The allocation is plausible and the large-sample cell-count check passes. Randomization, exclusions, multiplicity, and outcome integrity still depend on the experiment design.
Design around the smallest effect worth acting on.
Equal-allocation approximation using the observed control rate, 95% two-sided confidence, 80% power, and a 8% relative change.
This is a normal-approximation planning estimate, not a guarantee. Increase it for attrition, clustering, repeated looks, multiple variants, or conservative operational buffers.
The observed lift is one estimate, not the answer.
Two randomized samples would rarely produce identical rates even when the underlying experiences perform the same. The interval shows a range of effect sizes compatible with the design and observed data under the method’s assumptions. A wide range means the experiment has not measured the magnitude precisely.
Still compare the range with the effect needed to justify action.
The data remain compatible with some benefit and some harm.
For a higher-is-better metric, B appears worse.
A detectable effect may still be too small to matter.
Define the minimum worthwhile change before reading results. Include the cost of implementation, downstream quality, customer trust, revenue, complaints, and reversibility. A tiny precisely measured increase can be statistically distinguishable while remaining operationally irrelevant.
A p-value does not tell you the probability that B is better, the size of the effect, or whether the change is worth shipping.
An unexpected allocation can invalidate the comparison.
A sample-ratio mismatch test asks whether the observed A/B allocation is unusually far from the planned split. A strong mismatch can indicate assignment, filtering, delivery, eligibility, logging, bot, or identity-resolution problems. Investigate the pipeline before interpreting conversion differences.
Repeatedly checking and stopping at p<0.05 changes the error rate.
A fixed-horizon test chooses the sample, duration, metric, exclusions, and analysis before launch, then reads the primary result at the planned end. If the team needs continuous monitoring or early stopping, use a valid sequential design with its own boundaries. Keep campaigns running through the business cycle needed to represent weekday, time-zone, and delayed-conversion behavior.
Before declaring a winner.
Is p=0.03 a 97% chance that B is better?+
No. A frequentist p-value is calculated under a no-difference model; it is not the posterior probability that a variant wins or that the null hypothesis is true.
Does an inconclusive result prove the variants are equal?+
No. It means the data did not distinguish the effect from zero at the selected interval level. Read the interval to see how much benefit and harm remain plausible. Equivalence requires a predefined margin and an appropriate equivalence design.
Should I test open rate?+
Open events can be generated, blocked, or obscured by mailbox privacy and image-loading behavior. If opens are used, treat them as client-dependent signals and pair them with downstream clicks, conversions, complaints, or other customer outcomes.
Can I count every email send as an independent unit?+
Only when assignment and outcome observations are genuinely independent at that level. Repeated sends to the same recipient, account, or household create correlated observations that a simple two-proportion calculation does not model.
What if I tested several subjects or metrics?+
Selecting the best result across many variants, metrics, segments, or repeated looks increases false-positive risk. Predefine one primary comparison or use a method that accounts for multiplicity.
Run email experiments in context.
Connect campaigns, customer activity, delivery signals, and downstream outcomes in one workspace.
Start for free