Most A/B testing advice is written for stores with traffic to burn, then handed to merchants doing a few thousand sessions a month as if the maths didn’t change. It does. Low-traffic stores can test, but only certain kinds of tests, run a certain way, and sometimes the most profitable decision is not to test at all. Here’s the honest version.

The uncomfortable maths first

A test needs enough visitors to separate a real effect from random noise, and the smaller the effect you’re chasing, the more visitors that takes, the requirement scales with one over the effect size squared. Using the standard two-proportion formula (95% significance, 80% power, two-tailed) with a 7% baseline add-to-cart rate:

Lift you want to detectVisitors needed (total)At 50/dayAt 150/dayAt 400/day
+10% relative (7% → 7.7%)~43,600~2.4 years~10 months~16 weeks
+25% relative (7% → 8.75%)~7,400~5 months~7 weeks~3 weeks
+40% relative (7% → 9.8%)~3,100~9 weeks~3 weeks2 weeks*

* Held at a two-week minimum regardless.

(“Visitors” means visitors to the pages actually being tested, not sitewide sessions, a distinction that quietly doubles or triples many merchants’ real timelines.) The full derivation and calculator inputs are in our guide to sample size and test duration, and you can run your own numbers in the sample size calculator.

Two blunt conclusions:

  1. At low traffic, small optimisations are untestable. A 10% lift, a genuinely good result, would take a 50-visitor-a-day store years to confirm. Nobody runs that test; anyone who claims to has actually run a shorter test and read noise.
  2. Big effects remain testable. A 40% swing at 150 visitors a day resolves in about three weeks. The game isn’t closed to small stores; the table stakes are just higher per bet.

Everything below follows from those two facts.

Strategy 1: test bigger swings

Since sample size falls with the square of the effect, the single best lever a small store has is to test changes big enough to plausibly move the number a lot. Button colours won’t. Candidates that might:

  • A restructured product page layout, media, buy box and social proof rearranged, not one element nudged
  • Completely different photography style (lifestyle vs studio) across the gallery
  • A long-form, benefit-led description versus a short spec sheet
  • Adding or removing whole sections: size guides, comparison tables, shipping/returns reassurance up top

The mindset shift: at low traffic you’re not sanding edges, you’re choosing between genuinely different pages. Our list of product page A/B test ideas is worth filtering through exactly one question, “could this believably change add-to-cart rate by a quarter or more?” If not, don’t spend your traffic on it.

Strategy 2: put tests where the traffic already is

If one product gets 200 visits a day and your long tail gets 5 each, test on the bestseller. A significant result on the page that earns most of your revenue beats an eternally inconclusive result somewhere unimportant. Check your landing-page report before designing any test, and be honest about where visitors actually arrive.

The same logic applies to metrics: measure the most frequent meaningful event. Add-to-cart happens several times more often than purchase, so an add-to-cart test reaches a verdict several times sooner, one reason it’s the sensible primary metric for product page tests (the trade-offs are covered in the complete guide to Shopify A/B testing).

Strategy 3: pool products that share a template

Here’s the most under-used option on Shopify. Your products don’t each have a bespoke page, they share templates. Test at the template level and every visitor to every product on that template feeds the same experiment.

Twenty products averaging 15 daily visitors each are individually untestable. Pooled on one template, that’s 300 visitors a day, enough to detect a 25% lift in roughly three and a half weeks. This is precisely how Atchoo! works: you create an alternate product template in your existing theme (a Shopify template suffix like product.variant-b, no code, no theme duplication), and Atchoo splits visitors 50/50 across every product using it, tracking add-to-cart first-party and reporting a Bayesian probability that the variant wins.

One honest caveat: a pooled test answers “which template works better on average across these products”. If the pool mixes wildly different product types, a template that helps one category can mask harm to another. Pool products that sell in similar ways.

Strategy 4: run longer windows, within reason

If the maths says five weeks, run five weeks. Full-week blocks, start and end on the same weekday, minimum two weeks. Low-traffic stores should also expect lumpier data: one weekend flurry or one bulk order distorts a small sample far more than a large one, which is another argument for longer windows.

But there’s a ceiling. Past six to eight weeks, cleared cookies, device switching and seasonal drift erode the experiment’s integrity, and slow tests carry a real opportunity cost, that’s two months you couldn’t test anything else. If a test needs a quarter of a year, the correct fix is a bolder variant or a bigger pool, not a longer calendar.

Strategy 5: sequential learning over one-off verdicts

High-traffic teams can treat each test as a standalone trial. Low-traffic stores do better treating testing as a compounding research programme:

  • Keep a test log, hypothesis, dates, numbers, verdict, and what you concluded. Ten “inconclusive” tests that all leaned the same direction are collectively telling you something no single test could.
  • Let results seed the next test. A weak-but-positive signal for longer descriptions becomes next month’s bolder version of the same idea.
  • Accept graded evidence deliberately. Where a giant retailer demands 95%+ before shipping, a small store can rationally act on an 85–90% Bayesian probability-to-win for cheap-to-reverse changes, provided you decide that threshold in advance and log it, rather than sliding the bar to wherever today’s dashboard sits. What “probability to win” means, and why it degrades more gracefully than p-values at modest samples, is covered in Bayesian vs frequentist A/B testing.

Lower thresholds mean more false positives; that’s the honest price. The discipline is accepting it knowingly, not pretending underpowered tests are conclusive.

When NOT to A/B test

Sometimes the best testing strategy is none. Skip A/B testing when:

  • Under ~1,000 monthly visitors to your product pages. Even huge effects take months to confirm. Your growth constraint is traffic, not conversion, spend the effort on acquisition.
  • Best practice needs no referendum. Fixing broken mobile layouts, illegible text, missing product information or a three-second-slower page isn’t a hypothesis. Ship it.
  • The change is reversible and low-risk. Ship it, watch your baseline for a few weeks, move on. Imperfect before/after observation beats a test that was never going to conclude.
  • You can get richer answers cheaper. At small scale, five recorded user-testing sessions or a dozen customer conversations often teach you more than a quarter of inconclusive split tests. Use qualitative research to find the big swings, then A/B test only the ones worth confirming.

A/B testing earns its keep when a change is genuinely uncertain, plausibly large, and expensive to get wrong. That bar is higher at low traffic, which mostly means fewer, better tests.

Frequently asked questions

What’s the minimum traffic for A/B testing on Shopify?

As a rough practitioner’s line: below about 1,000 monthly product-page visitors, don’t bother; from roughly 5,000 a month (~150/day to tested pages), bold changes become testable in 3–7 weeks; above ~20,000 a month you can start detecting more modest lifts. Your baseline rate moves these lines, so run your own numbers in the sample size calculator.

Can I just run my test for six months to compensate?

Not usefully. Beyond six to eight weeks, cookie churn and seasonal drift degrade the sample, and the opportunity cost compounds. Increase the effect size or pool more traffic instead.

Is it OK to call a winner at 85% probability?

It can be a rational choice for low-risk, easily reversed changes, if you set that threshold before the test and accept that more of your “winners” will be flukes. It is not OK to decide 85% is fine only after seeing 85% on the screen.

Does Atchoo work for low-traffic stores?

Yes, with the same honest caveat as everything above: it can’t conjure significance from traffic that isn’t there. Because Atchoo tests at the template level, it naturally pools visitors across every product sharing a template, the most effective sample-size lever a small store has, and its Bayesian readout gives you usable graded evidence rather than a pass/fail cliff. There’s a 14-day free trial if you want to see how quickly your traffic can support a verdict.