What is A/B testing?
A/B testing splits live traffic between two versions of something and measures which produces more of the outcome you care about.
What you have now
The existing page, flow or message, left alone. It is the benchmark, and it runs at the same time as the alternative rather than before it.
The one change
Identical except for the thing you are testing — ideally one meaningful difference, if the goal is to isolate what caused the result.
Visitors are assigned to one side or the other at random, which is what makes the two groups comparable.
Both run simultaneously, so time-based effects — a promotion, a season, an unusual traffic week — are less likely to bias one version over the other.
One success metric decided in advance — conversion, revenue per visitor, add-to-cart — with secondary and guardrail metrics alongside it if needed. Picking the primary one afterwards is how a test gets talked into a result.
The anatomy of a test. Everything above exists to make one comparison trustworthy.
That is the whole idea, and the randomisation is what earns it. With proper randomisation, simultaneous exposure and enough data, a statistically credible difference can be attributed to the change rather than to obvious time-based confounders like a season or a campaign.
The short version. A/B testing is not change something and watch the numbers. It is a comparison designed so the change is the most credible explanation for a statistically reliable difference. Everything else — the random split, the shared time window, the metric chosen in advance — exists to protect that one claim.
What is statistical significance in A/B testing?
Two versions of a page will always produce slightly different numbers. Significance is how you judge whether the observed gap is distinguishable from ordinary statistical noise.
It indicates that the observed difference would be unlikely if the two versions truly performed the same. Larger samples and larger effects make that easier to establish, and it is only half the question — the effect size and its confidence interval tell you whether a real difference is also big enough to act on. Tools handle the calculation. One buyer describes theirs at its most practical: it tells when we hit statistical significance, and shows green once an experiment's stats are positive or it tells you to roll it back if it's negative.
Buyers who rate their platform highly often name this as the reason. One gives a top score to a testing tool because the confidence intervals, the statistical significance, everything else built in there is really good.
How much traffic do you actually need?
More than most people expect — and this is the constraint that decides whether testing is open to you at all.
The arithmetic is unforgiving in one specific way: the smaller the lift you want to detect — your minimum detectable effect — the more traffic the test usually needs. Spotting a large difference takes relatively few visitors. Spotting a one or two percent lift can take a great many, because the effect has to stand clear of ordinary day-to-day variation.
Traffic and minimum detectable effect are the two levers people notice, but they are not the whole calculation. Required sample size also depends on your baseline conversion rate and on the statistical power and significance thresholds you choose — which together set the trade-off between missing a real effect and falsely concluding that one exists.
That is not a tooling problem, and no platform solves it. Buyers on lower-traffic sites run into it directly. One, running an information-led site rather than a store, puts it plainly: with really small numbers that come through, in order to reach statistical significance, it takes a long time even with A/B testing — so it is kind of limited for us in how broad the program can be.
The mirror image also shows up. Another buyer describes having the traffic and not using it: they are starting to get a lot of statistical significance across a lot of our different brands and products, and we're not really testing anything, which is a huge missed opportunity.
Check this before you shortlist anything. Traffic and the size of the change you want to detect together decide whether a test can finish in a useful timeframe. If they do not, the honest options are to prioritise higher-impact hypotheses, test on higher-traffic surfaces, or accept that the decision will be made on judgement rather than evidence — our reading, not a practice buyers describe.
What buyers call it
“Experimentation” appears far less often in buyer language than “A/B testing”.
Across these interviews, A/B testing is said roughly five times as often as experimentation — and about three-quarters of the buyers who do say experimentation say A/B testing in the same conversation. The second word rarely arrives on its own.
Where it does arrive, it tends to travel with a different kind of product. Buyers using the word experimentation are several times more likely to be running a developer-owned platform — Statsig, LaunchDarkly, Eppo — than buyers who only ever say A/B testing, though the absolute numbers on both sides are small.
That makes vocabulary a usable signal rather than a rule. In these interviews, experimentation is more strongly associated with developer-owned platforms and server-side testing, while A/B testing is more common in marketer-led page testing. Which of those you need is the fork this channel covers in detail.
A/B testing, multivariate, and feature flags
Three things that get used interchangeably and should not be.
Two versions compared against each other, often designed around one primary change when the goal is causal interpretation. This is what most people mean, and what most ecommerce traffic can actually support.
Several elements varied systematically, so you can estimate which changes — and sometimes which combinations — affect the outcome. Usually needs substantially more traffic, because the same audience is split across more variations. Useful when you genuinely need to test several factors at once, not a maturity step above A/B.
A switch that controls who sees a feature. Flags enable controlled rollout; experimentation adds randomised measurement on top. Because a flag already splits users, it can carry a test — which is why developer platforms show up in this category at all.
What do buyers rate the tools?
The spread is fairly flat: several enterprise and specialist tools sit in the same rating band.
| Tool | Rating | What it means here |
|---|---|---|
| Intelligems | 8.0 | The highest-rated tool here with a real sample, and a specialist in testing price and promotion rather than layout. |
| Statsig | around 8 | Too few buyers for a decimal. A developer-owned platform — the kind that travels with the word experimentation. |
| Shoplift | around 8 | Too few buyers for a decimal. A Shopify-native tool, and its buyers here are almost entirely small companies — averaging a few hundred employees. |
| AB Tasty | 7.5 | Rates above the two best-known enterprise names on a smaller but solid sample. |
| Dynamic Yield | 7.4 | Sits across testing and personalization, so buyers are scoring more than the test harness. |
| Optimizely | 7.3 | The most-rated tool in this set by a wide margin, and it rates in the same band as smaller specialist tools. |
| VWO | 7.3 | The other long-standing marketer tool, and the one this channel sets beside Optimizely. |
| Adobe Target | 7.3 | The enterprise suite option. One buyer scoring it top marks credits its confidence intervals and built-in statistics. |
| Monetate | 6.5 | The lowest-rated well-sampled tool here. |
| Kameleoon | not rated | Too few buyers rated it to publish a number, though it still appears on shortlists. |
The flatness is the point. Three of the best-known enterprise names land within a tenth of each other at 7.3, while smaller specialist tools sit above them. That fits what buyers describe elsewhere on this channel: what experimentation buyers wish they had known opens on the finding that the tool is rarely the bottleneck, with traffic, implementation capacity and testing discipline doing more to decide whether a program works.
Want this read against your own stack?
Get my read →Common questions
What is A/B testing?
A/B testing splits live traffic between two versions of something — the existing one and a changed one — and measures which produces more of the outcome you care about. The existing version is the control, the changed one is the variant, and visitors are assigned at random so the two groups are comparable. With proper randomisation, enough data and a valid test design, a statistically credible difference supports attributing the result to the variant rather than to obvious time-based differences between the groups. That attribution is the whole point, and it is what separates a test from simply changing something and watching the numbers.
What is statistical significance in A/B testing?
Statistical significance indicates that the observed difference would be unlikely if the two versions truly performed the same. Two versions of a page will always produce slightly different numbers, and significance is how you judge whether that gap is distinguishable from noise. It is only half the answer: the effect size and its confidence interval tell you whether a real difference is also large and precise enough to matter. Buyers describe their tools handling this for them — one says theirs simply shows green once an experiment's stats are positive, or tells you to roll it back if it is negative. The judgement it encodes is the reason testing is a discipline rather than a dashboard.
How much traffic do you need to run an A/B test?
There is no universal traffic number. Required sample size depends on your baseline conversion rate, the minimum detectable effect you care about, and the statistical power and significance thresholds you choose. The practical shape of it is that the smaller the lift you want to detect, the more traffic you need: spotting a large difference takes relatively few visitors, while detecting a one or two percent improvement can take a great many. This is the constraint that decides whether testing is available to you at all, and buyers on lower-traffic sites describe hitting it directly. One running an information-led site rather than a store says that with the small numbers coming through, reaching statistical significance takes a long time even with A/B testing, which limits how broadly they can test. No tool solves this; it is arithmetic.
What is the difference between A/B testing and multivariate testing?
An A/B test compares two versions directly, often designed around one primary change when the goal is to isolate its effect. A multivariate test varies several elements systematically so you can estimate which changes, and sometimes which combinations, affect the outcome. It usually requires substantially more traffic, because the same audience is split across more variations. It is not a maturity step above A/B testing — it is the right tool when you genuinely need to test several factors at once and have the traffic to support them.
Is A/B testing the same as experimentation?
They describe the same underlying method, but they are not used equally. Buyers in these interviews say A/B testing far more often than experimentation — by roughly five to one — and most of the people who do say experimentation also say A/B testing in the same conversation. Experimentation is the broader and more formal term, covering feature rollouts and server-side tests as well as page comparisons, and it travels more often with developer-owned platforms. If a vendor's site says experimentation and your team says A/B testing, that difference in vocabulary is itself a signal about which kind of product you are looking at.
How long should an A/B test run?
There is no universal number of days. A test should run until it has enough observations for the effect size you care about and has covered the meaningful traffic cycles for the business. For some ecommerce sites that means spanning both weekday and weekend behaviour, but there is no universal minimum number of days. Stopping the moment a dashboard first flashes significance can produce unstable conclusions, because early results can move sharply. Low-traffic sites, or tests chasing small expected lifts, may need considerably longer than a high-traffic test looking for a large effect.
What can you A/B test on an ecommerce site?
Any experience or rule that can be exposed consistently to randomised groups and tied to a measurable outcome — layout, copy, imagery, pricing and promotion presentation, checkout steps, search and recommendation behaviour. In these interviews, implementation depth often decides whether teams can test structural changes or only cosmetic ones. The questions worth asking a vendor are about that integration and about how results are reported, which the buyers who have already done this cover in detail.
Get this research made for your stack
Whether a testing program pays off depends on your traffic, your team and what you actually want to change. Do a 15-minute interview and get the read for your setup — what peers your size run, what they gave up on, and what it took to make it work.
Get my personalized read — 15-min interviewNo password · your interview is anonymized before it ever informs a page like this one.