The honest answer is “until you’ve collected enough visitors, and that number is knowable before you start.” For most Shopify stores that works out to somewhere between two and six weeks, but the range is wide, and picking a duration by feel is how tests end up producing confident-sounding nonsense. Here’s how to calculate yours properly.
Duration is a by-product of sample size
An A/B test doesn’t need a certain amount of time; it needs a certain amount of data. Time only enters the picture because data arrives at whatever rate your traffic delivers it. So the real question splits in two:
- How many visitors per variant do I need? (Pure maths, answerable in advance.)
- How long will my traffic take to supply them? (Division.)
Plus one constraint that isn’t about statistics at all: run in full-week blocks, for reasons we’ll get to.
The three things that drive sample size
1. Your baseline conversion rate
The rate your control currently achieves for the metric you’re testing. For product page tests this is usually add-to-cart rate rather than purchase conversion, it’s the metric the page most directly controls, and because it’s a more frequent event, it reaches significance sooner. (More on that trade-off in the complete guide to Shopify A/B testing.)
Rarer events need more data. Detecting a change in a 2% rate takes far more traffic than detecting the same relative change in an 8% rate, because each visitor carries less information.
2. Minimum detectable effect (MDE)
The smallest lift you’d care to detect. This is the single biggest lever on your sample size, and the relationship is brutal: halving the MDE roughly quadruples the traffic you need, because required sample size scales with one over the effect size squared.
An MDE is a choice, not a measurement. Ask: “what’s the smallest lift that would justify shipping and maintaining this change?” For most stores the honest answer is 10–20% relative, not 2%.
3. Statistical power
Power is the probability your test detects a real effect of the size you specified, if one exists. The standard choice is 80%: a real lift at your MDE will show up as significant four times out of five. Lower power means more real winners slip through undetected; higher power costs more traffic. Alongside power sits your significance level, conventionally 95%, as covered in our guide to statistical significance.
The actual maths, honestly
The standard two-proportion formula for the sample size per variant is:
n = (z_α/2 + z_β)² × [p₁(1−p₁) + p₂(1−p₂)] / (p₁ − p₂)²
where p₁ is your baseline rate, p₂ is the baseline lifted by your MDE, and the z-values are 1.96 (95% significance, two-tailed) and 0.84 (80% power), so the leading term is about 7.85.
Worked example: baseline add-to-cart rate 8%, and you want to detect a 20% relative lift (8.0% → 9.6%).
n = 7.85 × [0.08×0.92 + 0.096×0.904] / (0.016)²
= 7.85 × [0.0736 + 0.0868] / 0.000256
≈ 4,900 visitors per variant → ~9,800 total
Here’s that same calculation at an 8% baseline across a range of MDEs (95% significance, 80% power, two-tailed):
| Relative MDE | Variant rate | Per variant | Total visitors |
|---|---|---|---|
| +10% | 8.0% → 8.8% | ~18,900 | ~37,800 |
| +15% | 8.0% → 9.2% | ~8,600 | ~17,100 |
| +20% | 8.0% → 9.6% | ~4,900 | ~9,800 |
| +30% | 8.0% → 10.4% | ~2,300 | ~4,600 |
Note the quadratic pain: detecting a 10% lift costs roughly four times the traffic of detecting a 20% lift. You don’t need to run this formula by hand, our free sample size calculator does it, and estimates duration from your daily traffic.
What that means at different traffic levels
Divide total sample by daily visitors to the tested page, not sitewide sessions. If you’re testing one product template, only visitors landing on products using that template count. Assuming the 8% baseline table above:
| Daily visitors to tested page(s) | +10% MDE | +15% MDE | +20% MDE | +30% MDE |
|---|---|---|---|---|
| 200/day | ~27 weeks | ~12 weeks | ~7 weeks | ~3.5 weeks |
| 500/day | ~11 weeks | ~5 weeks | ~3 weeks | 2 weeks* |
| 2,000/day | ~3 weeks | ~9 days* | ~5 days* | 2 weeks* |
| 10,000/day | ~4 days* | 2 weeks* | 2 weeks* | 2 weeks* |
* Rounded up to a two-week minimum, never run shorter even when the maths says you could, for the reasons below.
Two honest conclusions jump out. First, high-traffic stores hit the sample long before the two-week floor matters. Second, at 200 visitors a day, chasing a 10% lift means a six-month test, which is why low-traffic stores need a different playbook built around bigger swings.
Why full-week cycles matter
Your traffic is not the same crowd every day. Weekend shoppers browse differently from lunch-break mobile visitors; email sends, ad schedules and paydays all inject different visitor mixes on different days. A test that runs Tuesday to Friday has only sampled Tuesday-to-Friday shoppers, and its result may not generalise to the week your winner actually has to survive.
So: always run in multiples of 7 days, and start/stop on the same weekday. If the maths says 17 days, run 21. A two-week minimum also dampens one-off distortions, a viral post, a competitor’s sale, a shipping delay, that can dominate a short test.
The flip side: avoid running much past six to eight weeks. Cookies get cleared, people switch devices, seasons shift, and your two groups slowly stop being clean samples. If a test needs three months to detect your MDE, the right fix is a bigger MDE (a bolder change), not a longer test.
Why stopping early corrupts results
The temptation is universal: day 5, the variant is up 22%, significance badge is green, why wait?
Because random noise wobbles, and early in a test the wobbles are huge. If you stand ready to stop the moment the dashboard flatters you, you will systematically harvest lucky wobbles and call them winners. Repeatedly checking with intent to stop, the “peeking problem”, can inflate your real false-positive rate from the nominal 5% to 25% or more. And even when an early-stopped winner is real, the measured lift is usually exaggerated, because you stopped at a peak. That’s how a “+22%” test turns into a change nobody can see in next quarter’s numbers.
The discipline: fix your sample size in advance, glance at the test only to confirm it’s healthy (traffic splitting evenly, events tracking), and judge it once, at the end. Tools with Bayesian readouts, Atchoo! reports a running probability that your variant template beats the control, hold up more gracefully to being watched, but “wait for your planned sample” remains the safest rule under any method.
A pre-launch checklist
- Pull your baseline add-to-cart rate for the pages under test (use at least 2–4 weeks of data).
- Choose the smallest lift worth shipping, be honest; 15–20% relative is a sensible floor for most stores.
- Run the numbers through the sample size calculator.
- Divide by daily traffic to the tested pages; round up to full weeks, minimum two.
- Write the plan down before launch. Then don’t touch it.
If your store runs on Shopify, Atchoo handles the mechanics, it splits traffic 50/50 between product page templates, tracks add-to-cart rates first-party, and applies Bayesian analysis to call the winner, with a 14-day free trial to run your first test properly.
Frequently asked questions
Is there a universal minimum duration?
Two full weeks. Even if your traffic delivers the required sample in three days, shorter tests over-represent whoever happened to be shopping that week and are vulnerable to one-off events. Two weeks captures two full weekday/weekend cycles.
Should I use sitewide traffic in the calculation?
No, only visitors who actually see the tested experience. For a product template test, that’s visitors to products using that template. This is the most common reason merchants underestimate duration.
What if I genuinely can’t reach the sample size?
Increase your MDE by testing bolder changes, test on higher-traffic pages, or pool products that share a template into one test. If none of that closes the gap, some changes are better shipped on judgement than tested badly, see our low-traffic testing guide.
Does the calculator work for revenue per visitor?
The two-proportion formula covers yes/no outcomes (added to cart or didn’t, purchased or didn’t). Continuous metrics like revenue per visitor need a different calculation and typically much more data, because a few large orders add enormous variance.