Conversion IntelligenceJuly 10, 20267 min read

You do not need more ideas. You need to know which one to test first.

Most teams have a backlog of thirty page improvements and capacity for four. Two thirds of tested ideas fail to move the metric they targeted. Choosing well matters more than testing more.

Share thisEmailLinkedIn

Most digital teams do not lack ideas. They lack conviction about which idea deserves to go first.

The list is always the same: rewrite the descriptions, add reviews, clarify shipping, fix the navigation, reshoot the imagery, surface the guarantee, rework the calls to action. Every item is plausible. Every item has an internal advocate. Capacity is four of them this quarter.

The base rate nobody plans around

Here is the uncomfortable part. Ronny Kohavi, who built the experimentation platform at Microsoft and later led work at Amazon and Airbnb, puts it plainly: over two thirds of ideas fail to move the metrics they were designed to improve.

It gets starker at the top end. At Airbnb, 250 ideas were tested for search improvement and 20 succeeded. Over 90% failed. Those 20 were worth a 6% improvement in booking conversion and hundreds of millions of dollars, which is the point: the failures are the cost of finding the winners.

But a small team does not get 250 attempts. If you can run four experiments a quarter and two thirds fail by default, the quality of your selection is not a detail. It is most of your result.

More testing is a strategy for a team with unlimited traffic and time. Better selection is the strategy available to everyone else.

Repeated friction is signal. A single comment is noise.

This is where a simulation run earns its place, and where it is easy to misuse.

A run produces a lot of observations. If you treat all of them as findings, you have replaced a backlog of thirty guesses with a backlog of thirty observations, which is not progress. The useful discipline is to look for repetition across different buyers.

One buyer confused by the shipping module is a data point that might be about that buyer. Five buyers with different priorities, budgets, and reasons for visiting, all stopping at the same shipping module, is a page problem. The repetition is the signal, not the eloquence of any individual complaint.

That distinction also tells you what kind of fix you need. A hesitation that shows up across every buyer type is usually structural: something absent, buried, or contradictory. A hesitation concentrated in one type is usually a segment gap: the page serves most people and loses a specific group, often the most risk-sensitive or the least familiar with the category.

Separate the quick fixes from the strategic gaps

Two piles, and they should never compete for the same slot in a roadmap.

Quick fixes are clarity problems. The return window is not stated near the price. Shipping timing is a range with no cutoff date. The review count is styled too quietly to register. These are copy and placement changes, they ship in days, and they are usually the highest ratio of impact to effort available.

Baymard's checkout research estimates the average large ecommerce site could gain a 35.26% lift in conversion from better checkout design alone, and a large share of that is this kind of work rather than anything architectural.

Strategic gaps are positioning problems. Buyers cannot tell why this product differs from the cheaper one. The page speaks to a use case that is not the one most buyers arrive with. The product genuinely does not fit a segment you are paying to acquire. These need research, merchandising decisions, sometimes product changes. They matter more and they do not belong in the same sprint as a copy fix.

The failure mode is putting both on one list sorted by enthusiasm.

Write each finding as a hypothesis

A finding is an observation. A hypothesis is something you can be wrong about, which is what makes it testable.

The translation is mechanical. Take the friction, name the segment, state the expected direction:

  • Finding: Several buyers could not tell what the materials were without opening a tab.
  • Hypothesis: Surfacing material details above the fold will reduce hesitation for quality-focused buyers and lift add to cart on this page.

That version tells you what to build, who it is for, and what would count as it having worked. It also tells you what would count as it having failed, which is the part teams skip and then argue about afterwards.

Simulation prioritizes. Live traffic decides.

The last step matters and it is the one most easily overstated. A run is a way to choose what to test. It is not a substitute for the test.

The sequence that works: run buyers against the page, collect the friction that repeats, sort it into quick fixes and strategic gaps, write the top few as hypotheses, ship the quick fixes, and put real traffic behind the ones that carry a real bet. You are using simulation to raise the quality of what enters the funnel of experiments, in a world where two thirds of what enters it will not work.

That is a smaller claim than replacing experimentation, and a more useful one. Better inputs, same rigor.

Before the season

This is the moment in the year when selection matters most. Between now and peak traffic there is time for a handful of changes, and after that every change is a risk taken against the quarter that pays for the year.

eLLMo runs test buyers matched to your real customers against your page and returns a ranked list of what stops people from buying, ordered by how widely each blocker appears across the panel. That ranking is the input to the roadmap, not another dashboard. See a live run and bring back a shortlist rather than a backlog.

*Related Links: 1,000 Experiments Club: A Conversation With Ronny Kohavi (AB Tasty), 50 Cart Abandonment Rate Statistics (Baymard Institute).*

Share thisEmailLinkedIn

See this in action on your page

eLLMo runs test buyers against your product page and returns a ranked list of what stops people from buying.

More from the blog