Buyer PsychologyJuly 15, 20267 min read

Your customer data is what makes a simulated buyer resemble a real one

Research on synthetic respondents is clear about where they hold up and where they break. The difference is almost always specification, and the best specification input a brand has is its own customer data.

Share thisEmailLinkedIn

The interesting question about simulated buyers is not whether they work. It is what makes one useful and another one plausible nonsense.

The published research is less equivocal on this than the marketing around it. Work out of Stanford in 2024 found that a language model given a rich individual profile could reproduce that specific person's survey responses with roughly 83% to 86% reliability. Other research is blunt about the failure mode: when synthetic respondents stand in for a representative survey population, variance collapses and a large share of statistical relationships shift.

Both findings point at the same conclusion. The quality of a simulated person is dominated by the quality of the description they were built from.

A simulated buyer is only ever as specific as the data used to describe them. Thin input produces confident, useless output.

What thin looks like

Most persona work in commerce is thin in a particular way. It describes a demographic and calls it a person: women 25 to 44, urban, health conscious.

That is a segment, not a buyer. It says nothing about what this person is worried about at the moment they reach a product page, what would make them leave, or what evidence would settle the question. Run a simulation on that and you get output that sounds reasonable and predicts nothing, because there was no signal in the input.

What actually adds signal

A home fitness brand we ran test buyers for is enriching its customer profiles with data that changes the picture: household income band, discretionary spend, and whether there are children in the home.

Consider what each does to a decision about a piece of equipment with a membership attached.

  • Presence of children changes both the space calculation and the time calculation. It moves the objection from is this good to where does this live and when would I actually use it, and the page has to answer a question it was not written to answer.
  • Discretionary spend determines whether the recurring membership is background noise or a monthly decision that gets re-examined. It predicts subscription retention better than income alone.
  • Income band against product price sets how much justification the page has to supply. The same $2,000 machine is an easy purchase for one household and a considered one for another, and the considered buyer needs a cost-per-year framing the easy buyer skips past.

None of that is guesswork. It is first-party data the brand already holds, plus enrichment, turned into the specification that makes a simulated buyer behave like a real one.

The practical sequence

  • Start from your actual customers, not your target market. Who buys is frequently not who the brand imagines, and the gap between the two is often the most valuable thing in the exercise.
  • Add the fields that change decisions, not the ones that are easy to get. Household composition, replacement cycle, the problem that brought them in. Skip anything that does not plausibly alter what a buyer notices on a page.
  • Include the reason for the visit. The highest-signal field is why this person is on this page today. Replacing something broken, acting on a New Year intention, buying a gift, and researching for later are four different buyers who look identical in a demographic profile.
  • Keep it inside your privacy commitments. This is aggregate specification work, not individual targeting. Describe segments, not people.

Ahead of the season

Peak season traffic is not your usual traffic. It is heavier on first-time buyers, heavier on gift buyers, and heavier on mobile than any other period of the year. A persona set built only from your existing customer base will under-represent exactly the buyers who arrive in November.

Which is an argument for doing this now rather than in the middle of it. The data you need already exists in your customer records. Turning it into a set of buyers you can test a page against is a matter of weeks, not quarters.

eLLMo runs test buyers matched to your real customers against your product page and returns a ranked list of what stops people from buying. The matching is the part that decides whether the list is worth acting on. For the research on why buyer specification dominates outcomes, see human heterogeneity in agentic markets and preference heterogeneity and persona validity.

See a live run, or bring us your customer profile and we will build the panel before your season starts.

*Related Links: Synthetic respondent reliability findings are from Stanford University research published in 2024 on generative agent simulations of individual survey responses.*

Share thisEmailLinkedIn

See this in action on your page

eLLMo runs test buyers against your product page and returns a ranked list of what stops people from buying.

More from the blog