Conclusion first: almost every virtual try-on "lift" number you'll see compares shoppers who chose to try on (already high-intent) against those who didn't. That measures who was going to buy anyway, not what the tool changed. The only comparison that isolates the tool's real effect is a randomized holdout — and because online fashion converts at just ~2–3%, the honest sample floor to detect even a 20% lift is roughly 16,800 sessions per arm. That math is why clean, holdout-measured lift numbers are rare, and why any single-round "X% lift" claim with no holdout and no denominator should be read as marketing.
Lift is a causal claim: because we added try-on, conversion went from X to Y. A causal claim needs a counterfactual — the same shoppers, same week, same catalog, without the try-on — to subtract out. Everything that isn't that counterfactual is measuring something else and calling it lift.
| How it's measured | What it compares | What it actually measures | Clean read? |
|---|---|---|---|
| Pre/post install | This month (with) vs last month (without) | The tool + seasonality + traffic mix + every other change | No — confounded |
| Tried-on vs didn't | Shoppers who opened try-on vs those who didn't | Shopper intent (they self-selected), not the tool | No — selection bias |
| Vendor case-study % | One store, one number, no control group | A marketing figure with no counterfactual | No — no baseline |
| Randomized holdout | A random 90% with try-on vs a random 10% without | The tool's own incremental effect | Yes |
Because shoppers choose whether to try on, and the ones who do are already leaning toward buying. So when tried-on shoppers convert at 18% and everyone else at 2%, most of that gap was there before the try-on rendered a single pixel. That's textbook selection bias, and it inflates "lift" by a lot — often by an order of magnitude.
Worked example: a store does 40,000 sessions/month, converts 2.5% (1,000 orders at a $90 AOV = $90,000). Its vendor reports tried-on shoppers convert at 20% and pitches a "+700% lift." Those shoppers self-selected — the number is their intent, not the tool's effect. The honest read costs nothing: withhold try-on from a random 10% from day one and compare the two random groups.
Pre/post feels rigorous, but you changed the calendar at the same time you changed the store. Fashion conversion swings with season, promotions, ad mix, and payday cycles — install in November and compare to October, and you've measured the holidays as much as the tool. The fix is a concurrent control: a group living in the exact same weeks, so seasonality hits both equally and cancels out.
Detecting a real effect at a low base rate takes a lot of traffic, and online fashion converts at only about 1.5–3% of visitors (IRP Commerce). Using a standard two-proportion power calculation (95% confidence, 80% power), here's the per-arm floor for a store at a 2.5% baseline — you need it in both the try-on group and the holdout:
Halving the effect you want to catch roughly quadruples the traffic; a lower baseline raises the floor. Small stores don't lack lift — they lack the sample to see it in under a quarter. Pick the smallest lift worth acting on, size the test for that, and run until you hit the floor — not until the number looks good (stopping when significance first flickers green manufactures false wins).
Even a real holdout can lie if the plumbing is off. The failure modes are catalogued in the experimentation literature (Kohavi, Tang & Xu, 2020):
The zombie stats we fact-checked earlier — 36/40/48/64% return reductions with no study behind them — exist precisely because this discipline is rare. Do the holdout and you replace every borrowed percentage with a number that's actually yours.
We run an always-on 10% holdout at live stores and publish denominators — not a single blended "lift multiplier," because a clean store-wide lift number needs the sample floor above. At one live Ello store, measured in its own dashboards: 18.3% of try-on sessions purchased (21 of 115),
,150 attributed revenue over 30 days, ~$0.067 compute per try-on, with a 10% always-on holdout. That 18.3% is a tried-on cohort with its denominator shown — it is NOT a holdout-measured lift claim, because that cohort self-selected (our data; one store, results vary). That discipline is the pitch for Ello: 2D AI try-on on your existing product photo, covering clothing and accessories, measured against a real holdout instead of a flattering cohort. Verify for yourself — compare the Shopify try-on apps and read the real client numbers with their sample sizes attached.
With a randomized holdout: withhold try-on from a random 5–10% of shoppers, keep everyone in the same weeks so seasonality cancels, and compare conversion between the two random groups. Pre-register the metric and required sample size before you start, and don't stop early when significance first flickers.
Because shoppers choose whether to try on, and the ones who do are already closer to buying. Most of the conversion gap is pre-existing intent, not the tool's effect — selection bias that reliably overstates lift.
At a 2.5% baseline, detecting a 20% relative lift at 95% confidence and 80% power takes roughly 16,800 sessions per arm (~33,600 total); catching a 10% lift needs closer to 64,000 per arm. Lower baselines and smaller target effects push the number up fast.
No — a before/after comparison confounds the tool with everything else that changed on the calendar: season, promotions, ad mix. Use a concurrent holdout group living in the same weeks instead.