Compare variants

How it works

the same result at three sample sizes
200A 3.0% · B 4.5% · +50% liftp = 0.42 · NOISE
1,000A 3.0% · B 4.5% · +50% liftp = 0.08 · LEANING
2,500A 3.0% · B 4.5% · +50% liftp = 0.006 · REAL
Identical rates. Only the sample size changed. Most cold email tests stop at the first row.
01

Enter the counts

Sends and the outcome you care about — replies, positive replies, meetings, opens — for a control and up to three variants. Use raw counts, not percentages.

02

Compare each variant to control

A two-proportion z-test with a pooled standard error gives a two-sided p-value. The 95% confidence interval on the difference shows how wide the uncertainty still is.

03

Read the verdict

Significant at 95% is the bar for calling a winner. Between 80% and 95% is a lean worth extending. Below that, the test has not said anything yet.

04

Plan the next one

Sample-size mode takes your baseline rate, the smallest lift worth detecting, and power, and returns sends per variant and days at your daily volume.

Why most cold email tests prove nothing

Reply rates are small. At a 3% baseline, a 200-send variant expects six replies, and the natural swing between two identical variants is plus or minus four. A test that stops at 200 sends per arm and declares the 9-reply subject line the winner is reading noise, and the “winner” will regress on the next campaign — which is exactly the pattern people describe when they say A/B testing does not work for cold email. It works; the samples are just an order of magnitude too small. To detect a 30% relative lift on a 3% baseline with the usual 80% power, you need about 4,000 sends per variant. That is the honest number, and it is why the sample-size mode is on this page.

What the p-value and confidence mean

The test asks: if both variants really performed the same, how often would a gap this large appear by chance? A p-value of 0.04 means about four times in a hundred. “Confidence” is one minus that, so 96%. The customary bar is 95%, and it is a convention rather than a law — for a low-stakes subject-line choice, 90% is a reasonable place to act, as long as you know you will be wrong one time in ten. The confidence interval on the difference is the more useful number: if it runs from −0.5 to +3.5 percentage points, the variant could still be slightly worse. When the interval excludes zero, the result is significant; the width tells you how much you learned.

Testing well

Split randomly, not by list or by inbox — variant A on the fresh domains and B on the old ones is a domain test. Change one thing. Run both arms at the same time; day of week and news cycles move reply rates more than most copy changes. Measure the outcome you sell on, which is usually positive replies or meetings, not opens — open tracking is unreliable since Apple Mail Privacy Protection and adds a pixel that hurts placement. And decide the sample size before you start; peeking at the results every day and stopping when B pulls ahead inflates the false-positive rate several-fold. The sequence planner and capacity calculator tell you how many sends a fleet can produce, which sets the ceiling on what is testable.

Multiple variants

Each variant is compared with the control separately. With three variants that is three tests, and the chance of at least one false positive at 95% rises to about 14%. The tool shows the Bonferroni-adjusted threshold alongside the raw one; if a variant only clears the raw bar, extend the test rather than ship it.

Frequently asked questions

Should I test opens or replies?

Replies, or positive replies, or meetings — whatever you actually sell on. Opens are inflated by privacy proxies and image-caching, and an open is not a result. Reply-rate tests need more sends, which is the honest cost of measuring something real.

What is a good minimum detectable effect?

Whatever lift would change your decision. A 10% relative lift on reply rate is rarely worth the sends to detect; 25–30% is the usual floor for copy tests. The calculator shows how quickly the required sample explodes as the effect shrinks.

Is a one-tailed test more appropriate?

Only if you would ignore a result showing the variant is worse, which you should not. The tool reports two-sided p-values; halve them if you insist on one-tailed.

Can I compare two campaigns run at different times?

You can enter the numbers, but the result mixes timing with copy. Anything not run concurrently on a random split is an observation, not a test.

Is any of this uploaded?

No. The maths runs in your browser.

Last reviewed

Related tools

What to run next

The checks that most often follow this one.

Planning & cost

More in this category

Read more

Guides that go deeper

Services

When the tools tell you something is wrong

The diagnostics here are free and always will be. When the fix is bigger than a DNS record, this is the work I do.

Get in touch

Start with a call

Bring a domain and the symptom. I will tell you what is actually wrong and whether you need me at all — plenty of people leave that call able to fix it themselves.

Thirty minutes, no pitch

We will run the checks together on your actual domains, and you will leave knowing what is broken, what it takes to fix, and what it should cost. If that is a job you can do in-house, I will say so.

Based inRangpur, Bangladesh — all time zones
RepliesWithin one business day
LicensingWorkspace below list price
Back to top