What should the primary metric in A/B testing be?
The metric gets chosen before the test, or it isn't a metric
The primary metric is the number that decides ship or don’t ship. One number. Agreed and written down before traffic starts flowing.
If it’s picked after the results come in, it isn’t a metric — it’s a justification. Any test tracks a dozen numbers, and in a dozen numbers something will always look good. A team that chooses its winner retroactively will ship a positive result roughly every time it tests, which is exactly the pattern you see in programmes that report a 70% win rate and flat revenue.
Everything else you measure is secondary: useful for understanding why a result happened, not for deciding whether it happened.
For ecommerce, the default is revenue per session
Revenue per session — total revenue divided by total sessions in each variant — is the metric that maps to the thing you actually care about. It moves when conversion rate moves, and it moves when average order value moves, which means it catches the trade that conversion rate hides.
That trade is common. Aggressive discount messaging, a simplified bundle, a lower-friction checkout that defaults to the cheapest shipping option: all of them tend to convert more people at lower basket values. Run the arithmetic on a variant that lifts conversion 6% and drops AOV 5% and you’re left with roughly +0.7% on revenue — inside the noise, and you’ve just shipped a permanent margin change on the strength of it.
The same logic applies on the other side. A variant that pushes a higher-value bundle can lose on conversion rate and still be the correct thing to ship. If conversion rate is your primary metric, you kill it.
When conversion rate is the right call
Revenue per session has a statistical cost. Order values are skewed — a handful of large baskets carry a disproportionate share of the total — so the metric is noisy, and noisy metrics need more traffic to reach the same confidence. In practice a revenue-per-session test often needs several times the sample of a conversion-rate test on the same page to detect an effect of comparable commercial size.
For a brand doing €1M–€5M online, that frequently means the honest answer is: you cannot power a revenue-per-session test on this page within a sensible timeframe.
When that’s the case, the call is conversion rate as the primary metric with AOV as a hard guardrail — not conversion rate on its own. The test wins if conversion rate improves and AOV has not degraded beyond a threshold you set in advance. You lose some sensitivity to genuine AOV upside, and you accept that consciously rather than by accident.
Two things worth doing before you settle for the fallback: run the power calculation properly, and consider trimming or capping extreme order values so a single outlier basket can’t decide your test.
Guardrails stop a win from becoming a loss
A guardrail is a metric that can veto a win. It doesn’t have to improve. It just isn’t allowed to get materially worse.
For most ecommerce tests the useful set is short:
- Average order value — mandatory whenever conversion rate is primary.
- Return rate or contribution margin — a variant that sells more of a heavily-returned category is a revenue win and a margin loss. Fashion and furniture brands live and die on this one.
- Page performance — heavy test variants that slow the page suppress the very behaviour you’re measuring.
- Downstream funnel steps — a change that lifts add-to-cart and drops checkout completion has moved friction, not removed it.
Engagement, time on site, and click-through rate belong here too, in the diagnostic column. They explain a result. They should never decide one. A test that improves click-through rate on a category tile and does nothing to revenue per session has not earned a deployment.
The metric only works if the data underneath it does
All of this assumes your revenue numbers are correct at the variant level, and in a lot of setups they aren’t.
The recurring problems are boring and expensive. Purchase revenue reported inclusive of VAT in one system and exclusive in another, so the same test reads differently depending on which report the team opens. Consent gating that suppresses a portion of conversions unevenly across variants. Duplicate purchase events inflating high-traffic segments. Test assignment that doesn’t survive the browser-to-server handoff, so users get counted in one variant and converted in the other.
Any of these will produce a test result that is internally consistent, statistically significant, and wrong. Before you argue about which metric to optimise, confirm the metric is being measured accurately in both variants.
What to do before the next test goes live
- Write down the primary metric, the guardrails, and the minimum effect worth shipping — before traffic starts.
- Run the power calculation on that metric. If the required sample exceeds what the page will see in six weeks, change the metric or change the test.
- Default to revenue per session. Fall back to conversion rate with an AOV guardrail only when traffic forces it, and note that you did.
- Verify that revenue is reported the same way, on the same basis, in the tool making the decision.
The point of a testing programme isn’t a win rate. It’s a set of shipped changes that show up in the revenue line. Choosing the metric before the test is what connects the two.
If your test results and your revenue reporting disagree, that’s a measurement problem, not an experimentation problem.
A Tracking & Data Audit establishes whether the numbers your tests are judged on are correct — 5 business days, €1,500 fixed.
