Guides6 min read
How to A/B test a 3D product page properly
Metrics, sample size, test duration and the traps that make 3D A/B tests lie to you. A practical guide to proving whether 3D moves your numbers.
By CharpstAR · 4 Oct 2023 · Updated 22 Sept 2026
Part of 3D models for e-commerce: the complete guide

Most 3D rollouts are never tested. The models go live, the numbers go up or down for a dozen reasons, and someone writes a slide claiming credit or quietly drops the subject. If you are spending money on 3D, it is worth knowing whether it works on your site, with your products, at your traffic level. Here is how to set that test up so the answer means something.
We build the models, so we have an interest in the answer being yes. That is exactly why we would rather you ran a real test than believed a case study.
Decide what you are testing before you touch anything
The three tests people conflate:
- Does having a 3D viewer on the page change conversion? Control gets the current page, variant gets the same page with a 3D viewer. This is the business question.
- Does interacting with the 3D viewer predict purchase? This is not a test. It is a correlation, and it will always look spectacular because people who engage deeply with a product page are people who were already close to buying.
- Which 3D placement works better? Above the fold versus in a tab, viewer-first versus image-first. Worth running, but only after the first question has a yes.
Almost every impressive 3D statistic you have read is the second kind. Shopify's widely quoted finding that products with 3D or AR content converted 94 percent more often compares product interactions across merchant data rather than running a controlled trial, and merchants model their best products first. The Rebecca Minkoff figures, 44 percent more likely to add to cart after viewing a 3D model and 27 percent more likely to order, are also conditioned on interaction. They are real and useful. They are not what your test will produce.
Pick the right primary metric
One primary metric. Everything else is secondary and cannot be used to declare a win.
For most furniture and home brands the right primary metric is conversion rate on the tested product pages, measured as orders containing a tested product divided by sessions that viewed a tested product page. Revenue per session is a reasonable alternative if your basket sizes vary a lot, but it is noisier and needs more traffic.
Bad primary metrics, and why:
- Time on page. A 3D viewer increases it mechanically. It proves nothing about sales.
- Bounce rate. Same problem, opposite direction.
- Add to cart alone. Fine as a secondary, but a 3D model that increases carts and not orders has moved the problem downstream, not solved it.
- Anything conditioned on interacting with the viewer. By definition only the engaged cohort can be in it.
Useful secondaries: add to cart rate, AR open rate on mobile, viewer engagement rate, and return rate on tested products over a longer window.
Work out whether you have enough traffic
This is where most 3D tests are doomed before they start, and it is arithmetic, not judgement.
A rough working rule for a two-sided test at conventional confidence and power: to detect a relative lift of about 10 percent on a 2 percent baseline conversion rate you need on the order of tens of thousands of sessions per variant. To detect a 5 percent relative lift you need roughly four times that, because required sample scales with the inverse square of the effect size. Halve the effect you want to detect and you quadruple the traffic you need.
Run the numbers with a sample size calculator using your own baseline conversion rate before you commit. If the answer is that you need six months, you have three honest options:
- Widen the test. Include more products in both arms so each arm accumulates sessions faster.
- Accept a bigger detectable effect. Decide up front that you are testing for a large lift, and treat a null result as "not large", not as "no effect".
- Do not run a conversion test at all. Measure engagement and returns instead, and make the decision on cost and customer feedback. That is a legitimate choice for a small shop.
What you must not do is run an underpowered test, see a positive swing, and stop.
Set the duration properly
Two rules, and the longer of the two wins.
At least two full weeks, ideally four. Weekday and weekend shopping behaviour differ, and a test that runs Tuesday to Friday measures Tuesday-to-Friday shoppers.
At least one full purchase cycle. Furniture is considered. People visit, leave, discuss it at home, and come back a week later. If your typical time from first visit to order is nine days, a ten-day test attributes purchases to a period that did not include the decision.
Fix the end date in advance. Peeking daily and stopping when the line looks good is the single most common way to manufacture a fake win.
The traps specific to 3D tests
Page speed contamination. If the variant loads a heavy model badly, you are testing page weight, not 3D. Lazy-load the model, keep it near the roughly 4 megabytes Shopify suggests as a target, and confirm that the variant's largest contentful paint is not materially worse than the control's before you start.
Modelling only the good products. If you build 3D for your ten best products and compare them with the rest of the catalogue, you have measured which products are best. Randomise at the visitor level on the same set of products, not by choosing which products get models.
Uneven model quality. One badly proportioned model with the wrong fabric colour will drag the variant down and you will conclude that 3D does not work. Review every model against the physical product before launch.
Seasonality and promotions. A campaign that runs mid-test hits both arms if you have randomised properly. It does not hit both arms if you rolled 3D out to one country or one device type. Randomise by visitor.
No testing tool. Google Optimize and Optimize 360 stopped being available on 30 September 2023, and Google pointed people towards third-party providers instead. If your last test was in 2022, check what you actually have before planning anything.
Measure returns separately and slowly
Returns are the other half of the business case and they run on a different clock. You need the full return window plus processing, so a quarter at minimum. Track return rate on tested products and, if you have reason codes, the share of returns coded as size or expectation, because that is the slice a scaled 3D model can plausibly affect. The published benchmark worth citing is modest: Shopify reported Gunner Kennels reducing return rates by 5 percent after adding AR.
What a result actually licenses you to do
A clean positive on ten products tells you that 3D works for products like those ten. It does not tell you that it works for your accessories, your small items or your commodity lines. Roll out in the direction the test pointed, then test again at the edges.
Our clients tend to reach the same conclusion by category rather than across the board. Sweef, Contura and MELIMELI all landed on 3D for the complex, variation-heavy, physically large products where a photograph leaves the main question unanswered.
If you need assets to test with, we will build a 3D model of one of your products for free. Send us one product and it is yours to keep, which is enough to get a pilot page built. Pricing for a proper rollout is on prices, starting at 10 dollars per product per month on Basic with a 12-month term and a 100-variant minimum, and what we build is described on solutions.
Sources
- Shopify, 2020 merchant data on 3D and AR, AR shopping and 3D e-commerce.
- Google, Optimize and Optimize 360 sunset.




