Your A/B test did not have enough traffic to tell you anything
A meta-analysis of 115 real tests found about 70% could not detect the size of win they were looking for. That is worse than running no test at all, because a blind result still feels like evidence.
Here is a story every growth team has lived. You ship a test on the product page. You wait three weeks. The readout comes back at 2.4% against 2.5%, no significance, call it a draw. Somebody says the new version looks better anyway, so you keep it. Somebody else says the old one was fine, so you roll it back. Either way the decision gets made in about four minutes, by taste, and everyone writes down that it was tested.
The test did not fail. The test was never able to answer the question. And the three weeks bought you something worse than nothing, because now the guess has a number attached to it.
The math almost nobody runs first
Sample size is not a matter of opinion. Once you fix your conversion rate, the size of the win you are hunting, and how sure you want to be, the traffic you need is arithmetic.
Take a store converting at 2.5%. That is comfortably better than average: Littledata's benchmark of 2,800 Shopify stores put the typical rate at 1.4%, with the top fifth above 3.2%. Hold to the usual settings of 95% confidence and 80% power, and here is what it takes to see a win, per variant:
- A 20% lift needs about 17,000 visitors.
- A 10% lift needs about 64,000.
- A 5% lift needs about 251,000.
- A 4% lift needs about 390,000.
Double each of those for the two arms together. A 4% improvement, on a 2.5% base, is a 780,000 visitor question. Plenty of good brands do not send that much traffic to a single product page in a year.
The wins you are actually hunting are small
The obvious reply is that nobody tests for 4%. Everyone is going for the big one. The trouble is that the record says 4% is roughly what a real win looks like.
Georgi Georgiev ran a statistical meta-analysis of 115 A/B tests published by GoodUI, using the traffic and conversion counts reported for each one. These are not hypotheticals. They are real tests that real teams ran and were proud enough of to publish.
Across the 85 tests left after he pruned the worst cases, the mean lift was 3.77% and the median was 3.92%. Only 31 of the original 115, about 27%, were statistically significant positives.
Now put the two halves together. Typical honest win: around 4%. Traffic needed to see a 4% win on a decent conversion rate: about 780,000 visitors. That gap is the whole problem, and it does not close by wanting it to.
Georgiev found the gap directly in the data. About 70% of the tests, 80 out of 115, did not have the statistical power to detect the effects they were looking for.
Most tests in the wild are not measuring your idea. They are measuring noise, and reporting it with two decimal places.
So "no difference" was not a finding
This is the part that costs the most, because it hides inside a result that looks clean.
When an underpowered test comes back flat, that is not evidence the two versions perform the same. It is evidence the test could not see. Those are completely different facts and they get written into the same Slack message.
Georgiev checked exactly this. Of the tests that came back without a significant result, only about a third were sensitive enough to rule out a change of 12% or more. For the other two thirds, a 12% swing could have been sitting right there, in either direction, and the test would have shrugged.
Twelve percent is not a rounding error. On most catalogs that is the difference between a channel that pays for itself and one that does not. Teams retire good ideas on readouts that could not have detected a win that size.
The winner's curse on the other side
The tests that do cross the line have their own problem. In Georgiev's set, the significant tests averaged a 6.78% lift against a 3.77% average across the board, close to double.
That is what you would expect when a test can only reach significance by catching a favorable bounce. The underpowered test that reports a winner is not just uncertain about whether the win is real. It is systematically overstating how big the win is. He flags a likely cause in the same data: teams watching results as they come in and stopping when the line looks good.
So you forecast next quarter on 7%, build the roadmap around it, and get 3% if you get anything.
This is a false economy, and it is expensive
Name the pattern plainly. You paid for the testing tool. You paid an engineer to build the variant. You paid three weeks of calendar time, in a season where three weeks is most of the runway. Then you made the call on instinct anyway.
That is worse than skipping the test, for one reason. An honest guess stays open to argument. A guess wearing a p-value closes the conversation, gets repeated in the next planning meeting, and hardens into something the team believes it knows.
None of this is an argument against testing. Testing is how you settle a close call once you have the traffic to settle it. It is an argument against pretending the traffic is there.
Three moves that respect the arithmetic
- Work out the traffic before you build the variant. Ten minutes with your real conversion rate tells you whether the test can answer the question. If it cannot, you have saved three weeks and learned the same amount.
- Stop shipping tests that hunt small wins. If your volume can only see a 20% swing, only test changes that could plausibly move things 20%. Button colors and headline tweaks are not that. A different offer, a restructured page, or a removed step might be.
- Retire the phrase "no difference." Make the readout say what the test could actually detect. "We could not rule out a 15% change in either direction" is honest and points at the real next step.
Where simulation fits, and where it does not
Buyer simulation is not a faster way to get a significant result. It is not estimating a lift at all, so traffic is not the gate. That is the whole point of it.
The question changes. Instead of asking how much a variant moves a rate across a population, you ask why a specific buyer stopped. We run test buyers matched to your real customers through the page and return a ranked list of what blocked the purchase, in their own words. A missing return policy. A claim with nothing behind it. A shipping cost that showed up too late. Those are reasons, not rates, and a reason does not need 780,000 visitors to be legible.
That output is a different tool for a different job. It tells you what to fix and for whom, before you commit budget. When you do have the volume to settle a close call between two versions, run the test. Just run one that can see.
The teams getting the most out of testing are not the ones testing most often. They are the ones who know, before they start, which questions their traffic can answer and which ones it cannot. See a live run on your own page.
*Related Links: Analysis of 115 A/B Tests (Analytics-Toolkit), Ecommerce Conversion Rate Benchmark (Littledata), 5 things A/B testing can't tell you, What buyer simulations reveal that analytics miss.*
See this in action on your page
eLLMo runs test buyers against your product page and returns a ranked list of what stops people from buying.
More from the blog