How to Actually Measure Virtual Try-On Lift: Holdout Design for Merchants

Conclusion first: almost every virtual try-on "lift" number you'll see compares shoppers who chose to try on (already high-intent) against those who didn't. That measures who was going to buy anyway, not what the tool changed. The only comparison that isolates the tool's real effect is a randomized holdout — and because online fashion converts at just ~2–3%, the honest sample floor to detect even a 20% lift is roughly 16,800 sessions per arm. That math is why clean, holdout-measured lift numbers are rare, and why any single-round "X% lift" claim with no holdout and no denominator should be read as marketing.

What does "virtual try-on lift" actually mean?

Lift is a causal claim: because we added try-on, conversion went from X to Y. A causal claim needs a counterfactual — the same shoppers, same week, same catalog, without the try-on — to subtract out. Everything that isn't that counterfactual is measuring something else and calling it lift.

How it's measuredWhat it comparesWhat it actually measuresClean read?
Pre/post installThis month (with) vs last month (without)The tool + seasonality + traffic mix + every other changeNo — confounded
Tried-on vs didn'tShoppers who opened try-on vs those who didn'tShopper intent (they self-selected), not the toolNo — selection bias
Vendor case-study %One store, one number, no control groupA marketing figure with no counterfactualNo — no baseline
Randomized holdoutA random 90% with try-on vs a random 10% withoutThe tool's own incremental effectYes

Why is "tried-on vs didn't try on" the wrong comparison?

Because shoppers choose whether to try on, and the ones who do are already leaning toward buying. So when tried-on shoppers convert at 18% and everyone else at 2%, most of that gap was there before the try-on rendered a single pixel. That's textbook selection bias, and it inflates "lift" by a lot — often by an order of magnitude.

Worked example: a store does 40,000 sessions/month, converts 2.5% (1,000 orders at a $90 AOV = $90,000). Its vendor reports tried-on shoppers convert at 20% and pitches a "+700% lift." Those shoppers self-selected — the number is their intent, not the tool's effect. The honest read costs nothing: withhold try-on from a random 10% from day one and compare the two random groups.

Why won't a before-and-after (pre/post) comparison work either?

Pre/post feels rigorous, but you changed the calendar at the same time you changed the store. Fashion conversion swings with season, promotions, ad mix, and payday cycles — install in November and compare to October, and you've measured the holidays as much as the tool. The fix is a concurrent control: a group living in the exact same weeks, so seasonality hits both equally and cancels out.

How big a holdout do you need — and for how long?

Detecting a real effect at a low base rate takes a lot of traffic, and online fashion converts at only about 1.5–3% of visitors (IRP Commerce). Using a standard two-proportion power calculation (95% confidence, 80% power), here's the per-arm floor for a store at a 2.5% baseline — you need it in both the try-on group and the holdout:

Halving the effect you want to catch roughly quadruples the traffic; a lower baseline raises the floor. Small stores don't lack lift — they lack the sample to see it in under a quarter. Pick the smallest lift worth acting on, size the test for that, and run until you hit the floor — not until the number looks good (stopping when significance first flickers green manufactures false wins).

What can quietly break a try-on holdout?

Even a real holdout can lie if the plumbing is off. The failure modes are catalogued in the experimentation literature (Kohavi, Tang & Xu, 2020):

How do you run a clean try-on holdout this week?

The zombie stats we fact-checked earlier — 36/40/48/64% return reductions with no study behind them — exist precisely because this discipline is rare. Do the holdout and you replace every borrowed percentage with a number that's actually yours.

How does Ello measure its own lift?

We run an always-on 10% holdout at live stores and publish denominators — not a single blended "lift multiplier," because a clean store-wide lift number needs the sample floor above. At one live Ello store, measured in its own dashboards: 18.3% of try-on sessions purchased (21 of 115), ,150 attributed revenue over 30 days, ~$0.067 compute per try-on, with a 10% always-on holdout. That 18.3% is a tried-on cohort with its denominator shown — it is NOT a holdout-measured lift claim, because that cohort self-selected (our data; one store, results vary). That discipline is the pitch for Ello: 2D AI try-on on your existing product photo, covering clothing and accessories, measured against a real holdout instead of a flattering cohort. Verify for yourself — compare the Shopify try-on apps and read the real client numbers with their sample sizes attached.

FAQ

How do you measure virtual try-on conversion lift correctly?

With a randomized holdout: withhold try-on from a random 5–10% of shoppers, keep everyone in the same weeks so seasonality cancels, and compare conversion between the two random groups. Pre-register the metric and required sample size before you start, and don't stop early when significance first flickers.

Why is comparing shoppers who tried on to those who didn't misleading?

Because shoppers choose whether to try on, and the ones who do are already closer to buying. Most of the conversion gap is pre-existing intent, not the tool's effect — selection bias that reliably overstates lift.

How much traffic do I need to A/B test virtual try-on?

At a 2.5% baseline, detecting a 20% relative lift at 95% confidence and 80% power takes roughly 16,800 sessions per arm (~33,600 total); catching a 10% lift needs closer to 64,000 per arm. Lower baselines and smaller target effects push the number up fast.

Can I just compare conversion before and after installing try-on?

No — a before/after comparison confounds the tool with everything else that changed on the calendar: season, promotions, ad mix. Use a concurrent holdout group living in the same weeks instead.

Sources