Ecommerce Conversion Rate Optimization: A Funnel-Stage Playbook
David SertillangeIndependent experimentation specialistTL;DR
- →Plan optimisation by funnel stage — discovery, product, cart, checkout, post-purchase — and size each test by the revenue behind the leak.
- →Avoid the ecommerce-specific measurement traps: skewed revenue, seasonal baselines, guardrails on margin, and many tests running at once.
- →Ship the test on Shopify, Adobe Commerce or WordPress without flicker, and with a purchase event that actually carries the order value.
Ecommerce conversion rate optimization is the work of raising the share of visitors who buy, and the amount they spend when they do, by running experiments against the specific places a purchase falls apart. It is not a genre of tooling and it is not a checklist of best practices: it is a sequence of stages — discovery, product detail, cart, checkout, post-purchase — each with its own failure modes, its own tests and its own way of lying to you when you measure it. This page is organised that way because that is how the work is actually planned. A separate question is which software you need to run any of it — session recording, a testing platform, a feedback tool, an analytics warehouse — and that question has its own page on this site; nothing below is a product comparison. What follows is the funnel, and the statistics that keep an ecommerce result honest.
Two things make ecommerce different from optimising a SaaS signup or a lead form. The outcome variable is money, not a binary event, and money on a storefront is violently skewed. And the traffic arrives against a calendar — promotions, paydays, seasonal peaks — that moves conversion rate by more than most experiments do. Both of those are measurement problems, and both are covered below in Measuring ecommerce tests without fooling yourself, which is the section that turns the stage-by-stage list into something you can trust.
How ecommerce CRO is planned: by funnel stage
Every storefront has the same skeleton. A visitor arrives, finds a set of candidate products, evaluates one, adds it to a cart, pays, and either comes back or does not.
flowchart LR
A[Landing / category / search] --> B[Product detail page]
B --> C[Cart]
C --> D[Checkout]
D --> E[Post-purchase]
A -. exits .-> X[No purchase]
B -. exits .-> X
C -. abandons .-> X
D -. abandons .-> XThe value of the frame is that it turns a vague target — "increase conversion rate" — into a set of local questions with local answers. Site-wide conversion rate is a product of stage rates, so a 5% improvement at a stage that only 8% of visitors reach is worth much less than the same improvement one step earlier, and the arithmetic tells you that before you build anything. Pull the stage rates first: sessions reaching each stage, and the drop between consecutive stages. The largest absolute leak, weighted by the revenue behind it, is where the next test belongs.
The rest of this section is stage by stage. For each: what actually costs money there, what is worth testing, and the specific way that stage's measurement misleads.
Category and search: the discovery stage
This is where most of the traffic is and where most of it leaves, and it is the stage teams under-invest in because its wins do not look like conversion wins.
The failure modes are concrete. Search returns nothing for terms customers actually use — synonyms, misspellings, brand names you do not stock but should redirect. Default sort is "featured", which is often an artefact of an old merchandising decision rather than a relevance judgement. Filters do not carry the attributes people shop by — size availability, in-stock only, price bands that match how the catalogue is actually priced. Category pages show 60 products with no way to narrow them, so the visitor's only tool is the back button.
Worth testing: the default sort order; whether an in-stock filter is on by default; how many products are visible before a fold or a pagination break; whether zero-result searches fall back to a category rather than a dead end; whether the product card carries the one attribute that decides the click for your catalogue (price, rating, delivery date, size range).
The measurement trap here is that the natural metric — clicks into a product page — is easy to move and does not imply revenue. Sending more people into products they will not buy raises product-page traffic and lowers product-page conversion rate, and if you are watching that second number as a health metric you will misread a genuine win as a regression. Judge this stage on revenue per session for the whole visit, and treat downstream rates as diagnostics, not verdicts.
The product detail page
The product page is where a purchase is decided, and it is the stage where the gap between "what we say" and "what the buyer needs to know" costs the most.
The recurring failure is missing decision information. Not a shortage of copy — a shortage of the two or three facts this specific product's buyer needs before they will commit: fit against their body or their room, what the delivery date actually is for their postcode, what happens if it is wrong, whether the thing is in stock in their size at all. When those are absent the visitor leaves to find them elsewhere, and frequently buys elsewhere too.
Worth testing: the position and specificity of delivery and returns information relative to the buy button; image count and whether the first image shows the product in use or on white; how the size or variant selector behaves when a variant is out of stock (hidden, disabled, or offered with a back-in-stock capture); whether reviews are summarised where the decision is made rather than a scroll away; the presence and framing of a price anchor.
The trap at this stage is unit of analysis. A visitor sees many product pages in a session, so a per-pageview conversion rate counts one indecisive visitor many times and one decisive visitor once. Randomise and analyse by visitor, not by pageview, and report the per-visitor rate. Anything else quietly weights your result towards the browsing behaviour you were trying to change.
The cart
The cart is a short stage with an outsized abandonment rate, and most of what happens there was caused earlier.
The dominant failure is cost surprise: shipping, tax, duties or a fee appearing for the first time at the cart. The second is a cart that behaves like a filing cabinet — no way to edit quantity without a reload, no indication that stock is held or not held, a "continue shopping" path that discards the cart on return. The third is a discount field that teaches the visitor to leave and search for a code, which is a leak you built yourself.
Worth testing: showing the delivered total, including shipping, before the cart — on the product page or in a persistent mini-cart; the threshold and framing of free shipping; whether the promo-code field is collapsed behind a link; a saved-cart or email-my-cart affordance; the prominence of the primary checkout action against the secondary "keep shopping" one.
The trap is that cart tests move average order value in both directions at once, and the two effects cancel in the headline metric. A free-shipping threshold that raises AOV can lower order count, and revenue per visitor can end up flat while both underlying behaviours moved a long way. Report order rate and AOV alongside revenue per visitor, and decide in advance which of the three the test is allowed to trade away.
Checkout
Checkout is the highest-intent traffic on the site and the most expensive place to be wrong, because everyone here has already decided to buy.
The failure modes are old and well understood, and still ubiquitous: a forced account creation before payment; a form that asks for information the order does not need; address entry with no lookup; validation that fires on submit rather than on blur and throws the visitor back to the top of the form; a payment method your customers use that you do not offer; a mobile keyboard that does not switch to numeric for the card field. Each of these is a small friction with a measurable exit rate attached.
Worth testing: guest checkout as the default path; one page versus stepped, with visible progress; address autocomplete; wallet payments placed above the card form rather than beside it; error handling that preserves what was typed; whether trust and returns information sits inside the payment step rather than in the footer.
The trap here is the one that matters most in ecommerce: a checkout change can raise completion and destroy margin. Offering a payment method with a higher processing fee, or a free-delivery promise that costs more than the incremental orders earn, both look like clean wins on conversion rate. This is exactly what guardrail metrics are for — a small set of metrics you require not to move badly, declared before the test starts, checked with every readout. Contribution margin per visitor, return rate and delivery cost per order belong in that set on any checkout test. And a checkout stage is where fraud and chargeback rates deserve one too.
Post-purchase
The stage teams skip. It has no conversion rate of its own, and it is where the economics of the whole funnel are settled.
The failure modes: an order confirmation that is a receipt and nothing else; no delivery tracking, so the customer's next contact is a support ticket; returns treated as a cost centre and made deliberately awkward, which suppresses one return and loses a customer; no second-purchase prompt at the moment the product arrives, which is the highest-attention moment the relationship ever has.
Worth testing: the content of the confirmation page and the confirmation email — reorder, referral, account creation offered after the purchase rather than before it; proactive delivery updates; the returns flow's friction; the timing of the follow-up relative to the delivery date rather than the order date.
The trap is horizon. A post-purchase test is judged on repeat rate and on customer value over weeks, not on the session it happens in, so the experiment has to run long enough for the outcome to exist. Decide the observation window before you start — see how long to run an A/B test — and do not compare a treatment group's 30-day repeat rate against a control group that has only had 12 days to repeat.
Measuring ecommerce tests without fooling yourself
Everything above is a hypothesis generator. This section is what separates an ecommerce CRO programme from a list of ideas, and it is where a storefront's statistics differ most from a generic A/B testing checklist.
Revenue is skewed, and a naive t-test on it misleads
Conversion rate is a proportion and behaves itself. Revenue per visitor does not. Most visitors spend nothing, a majority of buyers spend near the median, and a very small number spend ten or a hundred times that. The distribution is zero-inflated and heavy-tailed, and a single B2B-sized order landing in one arm on the last day can flip the sign of the result.
The consequence is not that a t-test is forbidden — the central limit theorem still applies to the mean, and with enough orders the sampling distribution of the difference becomes approximately normal. The consequence is that "enough" is far larger than for a proportion, and that a handful of extreme values dominates the variance and therefore the width of your interval. The normality assumption in A/B testing works through what actually breaks and what to do instead: capping or winsorising the metric at a pre-declared percentile, analysing the decomposition (order rate × AOV) rather than the product alone, and using a bootstrap interval when the tail is doing the talking.
Whatever you choose, choose it before you look. A winsorisation threshold picked after seeing which arm the whale landed in is not an analysis, it is a decision.
-- Revenue per visitor, decomposed, with a pre-declared cap.
-- The cap value is fixed in the test plan before the test starts.
select
variation,
count(*) as visitors,
avg(case when orders > 0 then 1 else 0 end) as order_rate,
sum(revenue) / nullif(sum(orders), 0) as aov,
avg(revenue) as revenue_per_visitor,
avg(least(revenue, 500)) as revenue_per_visitor_capped
from experiment_visitors
where experiment = 'checkout-guest-default'
group by variation;
Reporting the capped and uncapped figures side by side is the honest presentation: if they disagree, the uncapped result is being carried by a few orders and the test has not finished telling you anything.
Seasonality and the promotional calendar
An ecommerce site's baseline conversion rate is not stationary. It moves with the day of the week, with paydays, with promotions you run and with promotions your competitors run, and around a seasonal peak it can double. Two consequences follow.
First, a test must cover whole business cycles. A result gathered over a Tuesday-to-Friday window is a result about Tuesday-to-Friday shoppers; run whole weeks, and prefer a number of weeks decided by the sample size and duration calculation rather than by when the number crossed a threshold.
Second, a test that straddles a promotion is measuring a different site in each half. Either exclude the promotional period from the analysis window — declared in advance — or accept that the estimate applies to a mixed population and say so.
The related distortion is the one that fires at the start of any visible change: regulars notice the new layout and click it because it is new, or bounce off it because it is not what they had memorised. Both effects decay. Novelty and primacy effects explains how to see them — by plotting the effect against exposure time and against new versus returning visitors — rather than shipping the first week's number.
Price is the special case. A price test changes the economics of the order rather than the likelihood of it, moves margin and revenue in opposite directions routinely, and carries fairness and legal constraints the rest of CRO does not. It has its own page: price testing.
Many stage-level tests at once
Once the funnel frame is working, the natural next move is to test in several stages simultaneously. That is usually right — it is how a programme gets throughput — but it introduces two distinct problems that are easy to conflate.
The first is interference: two tests on the same journey can interact, and a cart test running alongside a checkout test can produce a combination neither team designed. Running concurrent A/B tests covers when overlapping traffic is safe, when to isolate with mutually exclusive groups, and how to detect an interaction rather than assume there is none.
The second is multiplicity. Twenty tests at 5% significance produce about one false positive by construction, before anyone has looked at a segment. Add per-stage breakdowns and per-device splits and the number climbs fast. False discovery rate control is the practical answer for a programme running many tests: rather than demanding each test survive a stricter threshold, control the proportion of your shipped wins that are false, which is the quantity a business actually cares about.
Running the test on your stack
None of the above depends on a platform, but the mechanics of shipping the change do, and the two places a storefront experiment breaks are the same everywhere: the snippet firing after the page has painted, and the revenue event never reaching the experiment.
Flicker first. On a storefront the tested element is usually below the fold and rendered by a theme, so an anti-flicker approach that hides the whole page is worse than the flicker. Scope the hiding to the element being changed, and keep the experiment snippet synchronous and ahead of the theme's own scripts.
Revenue second. A conversion event that does not carry the order value gives you an order rate and nothing else, which throws away the metric this whole page is about. Send the value from the confirmation step, with the currency and the order id, and deduplicate on that id.
// Fire the revenue event from the order confirmation step, once per order.
window['optimizely'] = window['optimizely'] || []
window['optimizely'].push({
type: 'event',
eventName: 'purchase',
tags: {
revenue: Math.round(order.totalMinorUnits), // integer minor units, never a float
orderId: order.id,
currency: order.currency,
},
})
Where the snippet goes, and how the order value is read out of the platform, is the part that differs. The site has a walkthrough for each of the three common storefronts: Shopify, Adobe Commerce (Magento) and WordPress, each covering where the snippet is installed in that platform's template system and how to get a trustworthy purchase event out of its checkout.
Two platform-level cautions apply regardless. Hosted checkouts frequently sit on a different domain or an isolated context, so an experiment scoped to the storefront domain will not see the purchase unless the event is sent server-side or forwarded deliberately; check this before running any checkout test rather than after. And a theme update can move or remove the element a running experiment targets, so pin the selector to something stable and re-verify after any deployment.
Where to start
If the funnel has never been measured this way, do the arithmetic first: stage rates, drops, and the revenue behind each drop. Almost always the largest weighted leak is either discovery or checkout, and they demand different work — discovery is a merchandising and relevance problem, checkout is a friction problem.
Then run one test, on the largest leak, judged on revenue per visitor with a declared guardrail set and a duration decided in advance. A programme that ships one honest result a month beats one that ships four ambiguous ones, because the ambiguous ones get re-litigated and the honest ones compound.

Independent experimentation specialist
David Sertillange is an independent experimentation specialist with 10 years implementing Optimizely across enterprise programs. He specializes in Feature Experimentation, analytics integrations, and helping teams build a culture of data-driven decision making.
Related articles
Subscribe
Practical Optimizely tips, monthly. No fluff.