Most A/B tests don’t fail because the tool was bad or the maths was wrong. They fail because of a handful of completely avoidable human mistakes, the same ones, over and over, on stores of every size. The frustrating part is that a broken test doesn’t look broken: it produces confident-looking numbers that happen to be meaningless.

Here are the ten mistakes we see Shopify merchants make most often, why each one wrecks your results, and the fix for each. If you’re new to testing, read this alongside the complete guide to Shopify A/B testing, this list is the “what not to do” companion to it.

1. Stopping the test early because you’re winning

The most common mistake, and the most damaging. You launch a test, check it on day three, see the variant up 22%, and call it. What you’ve actually done is capture noise: small samples swing wildly, and if you peek daily and stop the first time you’re ahead, you’ll “win” a large share of tests where nothing changed at all. It’s the statistical equivalent of flipping a coin until you hit three heads in a row and declaring the coin biased.

The fix: set a stopping rule before launch, a target sample size or duration from a sample size calculation, and honour it. Run whole weeks (weekday and weekend traffic behave differently). Look at the dashboard as often as you like; just don’t act before the planned end.

2. Testing trivia

Button colour. A one-word headline tweak. 5px of padding. These tests feel safe, but tiny changes produce tiny effects, and tiny effects need enormous traffic to detect. On a typical Shopify store, a test capable of detecting a 2% relative lift would need more visitors than most merchants get in a year, so the test ends inconclusive, and everyone concludes “testing doesn’t work for us.”

The fix: test things big enough to matter, layout restructures, image strategy, how you present trust signals, description format. Our list of product page test ideas is deliberately weighted toward changes with plausible large effects. Rule of thumb: if you wouldn’t expect a visitor to notice the difference, your traffic won’t detect it either.

3. Testing without a hypothesis

“Let’s try a different layout and see what happens” is not a test, it’s a coin flip with extra steps. Without a hypothesis, you can’t design a focused variant, you can’t interpret the result, and you learn nothing transferable: even if the variant wins, you don’t know why, so you can’t apply the insight anywhere else.

The fix: write one falsifiable sentence before you build anything: “We believe [change] will increase [metric] because [reasoning about customer behaviour].” If you can’t fill in the “because,” you’re not ready to test, go collect evidence (session recordings, customer questions, reviews) until you can.

4. Too many variants for your traffic

Testing A vs B vs C vs D feels efficient, why not try everything at once? But every variant splits your traffic further: four variants means each one gets 25% of visitors, so reaching a verdict takes roughly twice as long as a two-way test, and comparing multiple variants against control multiplies your chances of a false winner. For most stores’ traffic levels, multi-variant tests are a trap.

The fix: run A vs B. Put your single strongest challenger against the current page. If you have three ideas, rank them by expected impact and test them in sequence, you’ll get more verdicts per year, not fewer. This matters double for low-traffic stores.

5. Changing things mid-test

The test is running and you spot a typo in the variant. Or marketing “just quickly” swaps the hero image on the control. Or you tweak the price of the tested product. Any of these mid-flight changes means the data before and after the change measured different things, and the averages blend them into something that measured nothing.

The fix: freeze both versions for the duration. Treat a running test like wet paint. If you find a genuine defect in a variant, fix it and restart the test from zero, painful, but honest. And keep a note of anything external that changed during the run (sale launched, influencer post went viral) so you can judge whether the window was representative.

6. Ignoring the device mix

You design and review both variants on your laptop. Your customers shop on phones. A variant that looks brilliant at 1440px can bury the add-to-cart button below three screens of content at 390px, and since mobile is typically 70%+ of Shopify traffic, the mobile experience is the test, whatever your desktop preview says.

The fix: design mobile-first, QA both variants on a real phone, and read the device breakdown when the test ends, as diagnosis, not as verdict (see mistake 7’s cousin in our results analysis guide). If a variant wins on desktop and loses on mobile, you’ve usually found a mobile layout bug, not a mysterious desktop preference. Our mobile product page guide covers what good looks like on small screens.

7. Redefining the metric after the fact

You declared add-to-cart rate the primary metric. The test ends flat on add-to-cart, but hey, time-on-page is up, and checkout rate looks better on Tuesdays, so really it’s a win, right? This is metric-shopping: scan enough numbers and one of them will always look good by chance. It converts every test into a “win” and your testing programme into a fiction.

The fix: name one primary metric before launch, for product page tests, usually add-to-cart rate, and let it deliver the verdict. Secondary metrics are context and hypothesis fuel, never a substitute verdict. Write the primary metric down where the team can see it, so the after-the-fact renegotiation can’t happen quietly.

8. Not QA-ing the variant

The variant template has a broken image on tablet, a JavaScript error that hides the reviews widget, or a buy button that doesn’t work with one specific product option combination. The test runs for four weeks and the variant “loses.” You didn’t learn that your idea was bad, you learned that broken pages convert badly, which you already knew. Worse, the result gets logged as a real learning and poisons future decisions.

The fix: before launch, walk through the variant like a customer: every device class, every product option, add to cart, through to checkout. On Shopify template tests this is easy, preview the alternate template with ?view= on the actual product URL and click everything. Check the browser console for errors. Ten minutes of QA protects four weeks of data.

9. Seasonality blindness

You run a test across Black Friday week and conclude your new urgency-heavy layout is a permanent winner. Or you test in the January lull and kill a variant that would have thrived in normal traffic. Discount-event shoppers, holiday gift buyers, and your regular customers are different populations, and a test measures whichever population happened to walk through the door.

The fix: know your calendar. Avoid starting or ending tests across major sales events, launches, or big campaign pushes unless the event itself is what you’re testing (we cover that special case in BFCM testing strategy). If a test must span unusual weeks, validate the winner in a normal period before treating it as settled.

10. Never shipping the learnings

The quietest mistake: the test ends, the winner sits in the tool, and nothing changes. The winning template never becomes the default. The insight (“our customers respond to shipping reassurance near the button”) never gets applied to the other 40 products. Nobody writes anything down, and eighteen months later someone proposes the exact same test.

The fix: every test ends with three actions. Ship it, make the winner the live template for everyone and remove the loser. Spread it, ask where else the underlying insight applies and queue those changes or tests. Log it, hypothesis, result, and one sentence on what you now believe about your customers. The log is the compounding asset; individual wins are just interest payments.

The pattern behind all ten

Notice that nearly every mistake is a discipline failure, not a knowledge failure: deciding things after seeing the data instead of before, changing things that should be frozen, skipping the boring QA and calendar checks. The merchants who get consistent value from testing aren’t statisticians, they’re the ones who write the hypothesis down, leave the test alone, and do something with the answer.

Tooling can enforce some of that discipline for you. Atchoo! runs product page template tests with deterministic visitor assignment (no mid-test bucket-switching), tracks add-to-cart as the primary metric by design, and its Bayesian results are readable without a statistics course, but no tool can stop you peeking at day three and calling it. That part’s yours. The Pro plan has a 14-day free trial if you want to put the discipline into practice on a real test.

Frequently asked questions

What’s the single most damaging mistake on this list?

Stopping early (mistake 1). It’s damaging precisely because it feels like winning, you ship a change backed by noise, log a false learning, and calibrate future expectations against a lift that never existed. If you fix only one habit, fix this one.

How many of these mistakes can a testing tool prevent?

Some. A good tool gives you deterministic assignment, a stable primary metric, sound statistics, and device/source breakdowns. It cannot write your hypothesis, stop you peeking, QA your variant, or make you ship the winner. Roughly half this list is irreducibly on you.

Is it a mistake to run more than one test at the same time?

On the same page, yes, don’t. On different pages (two different product templates, say), it’s generally fine for most stores: the interaction effects are usually small compared to the noise you’re already living with. Just don’t run overlapping tests whose changes could plausibly influence each other’s metrics.

I’ve made several of these mistakes already. Are my past results worthless?

Not worthless, but downgrade them: treat past “winners” from peeked-at or metric-shopped tests as hypotheses rather than facts. If one of them drove a big decision, rerun it properly, pre-registered metric, fixed duration, QA’d variants, and let the clean result stand.