Every A/B test needs one primary metric, the single number that decides the winner before you launch. For Shopify product page tests, the choice usually comes down to add-to-cart rate versus completed-purchase conversion rate, and the trade-off is real: one is faster and cleaner, the other is closer to the money.
This article walks through the case for each, why revenue per visitor is more dangerous than it looks, and how guardrail metrics let you get the best of both.
Why you only get one primary metric
The primary metric is the question your test answers. Everything about the test, the sample size, the duration, the decision you make at the end, is built around it.
If you declare three “primary” metrics and pick whichever one favours your preferred variant after the fact, you’ve dramatically increased your odds of a false positive. Statisticians call this multiple comparisons; practitioners call it cherry-picking. Either way, the fix is the same: choose one primary metric before launch, write it down, and use everything else as supporting context.
So which one?
The case for add-to-cart rate
Add-to-cart rate is the percentage of product page visitors who add the product to their cart. For a product page test, it has three big advantages.
1. More events means faster, more reliable answers
The statistical power of a test depends on how many conversion events you collect, not just how many visitors you get. On most stores, add-to-cart events outnumber purchases several times over, a store converting 2% of visitors to purchase might see 6–10% of product page visitors add to cart. Treat those figures as illustrative rather than targets; your own baseline is the only number that matters, and you can sanity-check it against typical add-to-cart rate ranges.
More events means you reach a defensible conclusion in a fraction of the time, or reach one at all, for lower-traffic stores, purchase-based tests on a single product page can take months to conclude, which usually means they never do.
2. It’s scoped to the thing you changed
You changed the product page, so it’s fair to measure the product page’s job: convincing the visitor this product is worth buying. Add-to-cart is the visitor’s direct response to that page.
Completed purchase, by contrast, sits at the end of a chain the product page doesn’t control: cart page, shipping rates revealed at checkout, payment options, discount code hunting, a checkout that’s slow on mobile. If a visitor loves your variant page, adds to cart, then abandons because shipping is $14.99, your variant did its job, but a purchase-based test scores it as a failure.
3. Less noise between exposure and measurement
Every step between “saw the variant” and “counted the event” adds noise: time delays (people buy days later), cross-device journeys (browse on phone, buy on laptop, often counted as two different visitors), and external interruptions. Add-to-cart usually happens in the same session, on the same device, minutes after exposure. Shorter chain, cleaner signal.
The honest weakness
Add-to-cart is a proxy. A variant could, in principle, inflate add-to-carts without producing more purchases, say, by hiding the shipping cost information that used to filter out non-buyers earlier. This is uncommon for typical layout, imagery and copy tests, but it’s why guardrails matter (more below).
The case for completed purchase
Purchase conversion rate, the share of visitors who complete checkout, is the truer metric. Nobody ever paid rent with add-to-carts.
Choose purchases as primary when:
- Your change plausibly affects the post-cart journey. If you’re testing anything about price presentation, bundles, or shipping expectations, the interesting effect may show up between cart and checkout, and add-to-cart could actively mislead you.
- You have serious traffic. If your product page gets tens of thousands of sessions a month, the noise argument weakens, you can afford the less efficient metric and measure closer to revenue.
- The stakes justify the wait. For a change you’ll roll out across your whole catalogue, waiting longer for a purchase-validated answer can be worth it.
And its honest weaknesses, stated plainly:
- Noisier and slower. Fewer events, longer delays, cross-device loss. Expect to need several times the test duration for the same statistical confidence, the maths behind this is covered in our piece on statistical significance.
- Influenced by factors your page can’t touch. Checkout speed, payment methods, shipping costs, even an unrelated email campaign that lands mid-test, all of it moves the purchase number without your variant being responsible.
- Attribution gets fuzzy. Which product page “caused” a five-item order? Assignment and credit rules start to matter, and different tools answer differently.
Revenue per visitor: tempting but treacherous
Revenue per visitor (RPV), total revenue divided by visitors in each arm, sounds like the ultimate metric. It captures conversion and order size. In practice it’s the hardest metric to test well, for one dominant reason: outlier orders.
Order values are heavily skewed. Most orders cluster near your average; occasionally someone buys ten of everything. Now imagine a $2,000 wholesale-style order lands in variant B by pure chance. In a test where each arm’s revenue is, say, $30,000, that single order is a 6–7% swing, bigger than most real effects you’re hoping to detect. Your “winner” is one customer’s shopping spree.
The consequences:
- You need much larger samples than for a rate-based metric, because the variance of revenue is so much higher.
- Results can flip late in a test when a whale order lands, which tempts people into exactly the kind of peeking-and-stopping behaviour that invalidates tests.
- Mitigations exist but add judgement calls. Capping or winsorising order values (e.g. treating any order above the 99th percentile as if it were at the 99th percentile) helps, but now your metric needs a footnote, and the cap choice can change the outcome.
RPV is a reasonable secondary metric to inspect, and a defensible primary for teams with statistical support and lots of traffic. For most merchants running product page tests, a rate metric, add-to-cart or purchase, is the sturdier choice.
Guardrail metrics: how to use a proxy safely
A guardrail metric is one you don’t optimise for, but check to make sure your winner isn’t winning by breaking something. Guardrails are how you use add-to-cart as primary without worrying about being fooled by it.
Sensible guardrails for a product page test:
| Guardrail | What it protects against |
|---|---|
| Purchases / downstream conversion | Variant inflating add-to-carts that never become orders |
| Bounce rate on the page | Variant confusing or annoying visitors |
| Average order value | Variant shifting people to cheaper options |
| Page load time | Variant assets slowing the page down |
The rule: your primary metric decides the winner; a clearly worse guardrail vetoes it. If add-to-cart is up 12% but purchases from those sessions look flat-to-down, don’t ship, investigate. Usually the story is benign (noise, or checkout friction unrelated to the test), but the check costs you nothing and occasionally saves you from a bad rollout. Interpreting these mixed pictures is covered in more depth in how to analyze A/B test results.
A practical decision framework
- Testing layout, images, copy, trust signals, CTAs on a product page? Primary: add-to-cart rate. Guardrails: purchases, bounce, AOV. This is the right call for the large majority of product page tests, especially under ~50k sessions/month.
- Testing something that changes buyer expectations about cost or delivery? Primary: purchase conversion (or run add-to-cart primary with purchases as a hard guardrail, if traffic won’t support a purchase-based test).
- Tempted by revenue per visitor? Use it as a secondary lens, not the decider, unless you have the traffic and the appetite to handle outliers properly.
This is also why Atchoo! is built around add-to-cart rate as the primary metric for its template tests: for page-scoped product page experiments it’s the closest controllable proxy to revenue, it concludes in weeks rather than months on typical Shopify traffic, and the Bayesian analysis on top gives you a plain-English probability that the variant is actually better. (Atchoo doesn’t do price or checkout testing, for those, purchase-level tooling is the right instrument.)
For the bigger picture on designing trustworthy tests end to end, see the complete guide to Shopify A/B testing. And if you want to see what an add-to-cart-focused template test looks like on your own store, Atchoo’s Pro plan has a 14-day free trial.
Frequently asked questions
Isn’t add-to-cart rate just vanity if people don’t buy?
It would be if add-to-cart and purchases were unrelated, but on a given store they’re tightly linked: the checkout funnel converts carts to orders at a fairly stable rate over the length of a test. When a variant page produces more carts and nothing else about the funnel changed, more orders follow. The guardrail check is there to catch the rare exceptions.
Can I just run the test until purchases are significant too?
You can try, but for many stores the required duration becomes impractical (multiple months per test), and long tests bring their own problems: seasonality drift, cookie expiry, and the temptation to peek. Run the numbers with a sample size calculator before committing.
What about click-through metrics like “clicked the size guide”?
Fine as diagnostics, dangerous as primaries. Micro-interactions are easy to move without moving revenue. Keep the primary metric as close to money as your traffic allows, add-to-cart is usually the nearest point with adequate signal.
Should mobile and desktop be judged separately?
Look at the split as a secondary analysis, but decide on the pooled primary metric unless you planned a device-specific test up front. Slicing after the fact is how false positives sneak in.