The question the research actually answers
Ask whether AI models accurately replicate human preferences and the answer is no — not in any direct, point-for-point way. An earlier model showed degenerate behavior in controlled intertemporal choice experiments: it almost always took the immediate reward regardless of how attractive the delayed option became. That is not a miscalibrated preference. That is a preference that does not exist in any recognizable human form. A newer model improved: it responded to interest rate changes in the correct direction, becoming more patient as the return on waiting rose. But its discount rates were dramatically above every human benchmark the research had produced — far more impatient than real buyers across every condition tested.
This is the finding Goli and Singh published from the University of Washington after running 138,600 experimental cells across 22 languages. Their conclusion was precise: directly eliciting preferences from frontier models yields misleading results. The absolute level of any preference signal is unreliable. That conclusion should have ended the conversation about AI simulation. It did not — because a different finding in the same data pointed toward something more useful.
The structure holds when the level does not
What the models got right was not how much buyers preferred an option. It was which buyers preferred more versus less — and by how much, relative to each other. When the researchers conditioned on detailed demographic and attitudinal backstories, the models produced outputs that were demographically coherent and correlated with real consumer panels. The heterogeneity — the shape of who differs from whom, and in which direction — was meaningful even when no individual output was calibrated to human levels.
This distinction is the theoretical foundation for what researchers call silicon sampling: using AI-generated personas as proxies not for individual humans but for the distribution of preferences across a population. The outputs are not predictions of what a specific buyer will do. They are estimates of how a buyer type responds relative to other buyer types — which segment is more price-sensitive, which hesitates at claims that others accept without friction.
For a simulation product, that reframe is the entire question. Which segments are being lost and at which page elements? That question requires only that the structure of differences is reliable. The research confirms it is.
Why persona conditioning beats direct elicitation
Direct elicitation — asking a model 'would you prefer $100 now or $120 in a month?' — produces the degenerate behavior documented above. The model has no context that would shape a preference, so it defaults to whatever pattern was most frequent in its training data. For marketing decisions, the result is similarly hollow: the model responds to a landing page as a generalized reader, not as any particular buyer.
Persona conditioning changes the computation. A detailed backstory — age, income bracket, purchase history, stated priorities, risk orientation, time horizon — gives the model a coherent identity to inhabit when processing the stimulus. The choice becomes a function of the persona rather than a function of training data frequency. The research confirmed that this shift produces outputs that track real demographic variation: older, lower-income personas showed greater patience in intertemporal choice, consistent with findings from human consumer research. The model was reproducing patterns from its training data about how different kinds of people behave, channeled through the persona description.
This is why eLLMo builds each persona around a calibrated backstory rather than a label. A persona is not a tag that says 'Skeptical Analyst.' It is a full identity — prior experiences with the category, what has disappointed the buyer before, what they are protecting — that produces a specific reading posture when the persona encounters a page.
Chain-of-thought: correction and confession
Asking a model to reason through a choice before answering does two things simultaneously. It moves absolute outputs closer to human behavior — the research found that chain-of-thought prompting reduced but did not eliminate the discount rate gap. And it exposes the factors the model actually weighted in producing the answer.
Goli and Singh applied LDA topic modeling to the reasoning traces and found structured patterns. As delay length increased, discussion of risk and uncertainty intensified. As interest rates rose, discussion of opportunity cost declined — the model recognized that a high enough return weakened the risk argument. Strong future-tense-reference language models discussed risk roughly 20% less than weak FTR models, tracking a documented pattern in human linguistic research.
The practical implication: the reasoning trace is not just an explanation for the verdict. It is a map of what factors a buyer in this condition is weighing. When a persona explains why it hesitated at a price anchor — 'the price appeared before any evidence of what it buys, so the number landed as a cost rather than a value' — the explanation names a specific, fixable problem. The chain-of-thought is not just a corrective mechanism. It is the product.
Language is not neutral substrate
The 22-language study produced a finding with direct operational consequences: the linguistic register of the prompt measurably shifts model outputs independent of semantic content. Models showed substantially more patience in weak future-tense-reference languages (German, Mandarin) than in strong future-tense-reference languages (English, Russian). Weak FTR languages do not grammatically separate the future from the present, reducing the felt distance between now and later. The models replicated this pattern across all conditions.
This means how copy is written affects how a persona reads it. A promise framed as future outcome ('You will see results in 30 days') activates a different reading posture than the same claim presented as present-tense outcome ('Customers see results in 30 days'). The first marks a temporal distance the buyer must bridge. The second presents the outcome as current reality the buyer is entering. eLLMo treats tense, voice, and register as testable simulation variables — running the same page with different copy framings against the same panel and reading the divergence in reasoning traces to identify which framing produces less hesitation for which buyer type.
A concrete buyer voice
Consider a high-Conscientiousness persona: 44-year-old operations manager, prior purchase of a similar product that underdelivered on stated specifications, current goal of finding something with verifiable performance claims. The persona encounters a page where the hero section leads with a bold outcome claim followed immediately by a CTA, with supporting evidence three scrolls down.
The reasoning trace: 'The headline makes a specific claim but does not explain the mechanism. The CTA appeared before I had enough information to act on it. I scrolled to find evidence and eventually found it, but the ordering signals that the brand prioritizes conversion over credibility.'
That is not a conversion prediction. It is a buyer thinking out loud about page architecture. The finding is not a percentage. It is: evidence-first buyers read this page as a trust deficit created by where the social proof sits relative to the claim. Specific. Fixable. Surfaced before a dollar of traffic ran.
How eLLMo operationalizes the research
Goli and Singh recommend chain-of-thought conjoint plus topic modeling as a hypothesis generation tool rather than a direct measurement instrument. eLLMo is the commercial application of that recommendation, applied to landing pages. A calibrated panel of personas — each built on a full demographic and attitudinal backstory, each conditioned to a specific buyer archetype — reads the page, produces a reasoning trace, and the traces are analyzed for recurring patterns: which friction themes appear across how many personas, at which elements, and with what intensity.
The result reads as buyers talking back to the page:
- 'The social proof section names clients but does not describe outcomes — high-skepticism buyers treat unnamed results as decoration.'
- 'The pricing section appears before the value stack is complete — analytical personas read this as a negotiation tactic.'
These findings are grounded in the research infrastructure: heterogeneity is reliable, reasoning traces expose real weighting factors, cross-persona pattern analysis is more durable than any individual absolute output. See how eLLMo maps buyer types and how simulation validity is established for the full architecture.
The commercial implication
Hypothesis generation is not a consolation prize for a tool that cannot predict. It is the correct instrument for the question teams face before a campaign fires.
A simulation product that claims to predict conversion rates makes a promise that a live campaign falsifies within days. A simulation product that surfaces friction patterns ranked by buyer segment makes a recommendation that is visible in the tests that do not need to run — because the friction was removed before the campaign started.
The research also points toward a durable calibration trajectory. Earlier models showed degenerate behavior; a newer model improved substantially. As frontier models improve at holding persona identity stable across a long reasoning trace, the gap between simulated and real buyer behavior narrows. A product systematically measuring that gap — running the same pages against simulation and then against real traffic, tracking which friction themes predicted real drop-off — builds a proprietary calibration dataset that compounds with every run. Goli and Singh identified the research gap. eLLMo is closing it.