Research Area 01

AI Buyer Behavior and Structural Bias

A substantial body of controlled research has tested how frontier AI models behave when placed in the role of autonomous buyers — making product selections, comparing options, and evaluating landing page content. The consistent finding is that AI buyers are economically rational in aggregate but exhibit systematic, non-human biases in how they process information.


Key takeaways

Position effects are structural, not visual

In controlled e-commerce experiments, AI agents exhibit strong position bias: items placed in certain structural positions within a product listing receive disproportionate selection probability. This bias persists even when product information is presented as plain text with no images or layout. It is not a visual processing artifact — it is how language models weight information in ordered sequences. On landing pages, this means content that appears earlier in the document's linear structure carries more weight in AI buyer reasoning, independent of visual hierarchy. Headlines and above-the-fold copy are not just attention-getters; they anchor the entire evaluation frame.

Framing and anchoring effects are measurable and exploitable

Replication studies of classic behavioral economics experiments confirm that AI agents exhibit the same framing effects documented in human consumer research: identical choices framed as losses ('avoid losing $10') produce different responses than the same choices framed as gains ('save $10'). Status quo framing — presenting one option as the established default — measurably increases its selection frequency, replicating the same status quo bias observed in human decision-making. Anchoring is also measurable: AI agents shown a reference price systematically weight that anchor in their valuation judgments, with anchor effects comparable in direction to human studies (though attenuated in magnitude). These effects are not incidental — they are encoded in how language models represent value from training on human-generated text.

Credibility signals vs. self-assertion: a persistent divide

Controlled experiments consistently find that AI buyers heavily weight third-party credibility signals — editorial curation, press mentions, analyst endorsements, verified reviews — and discount self-promotional language, superlatives, and uncorroborated claims. The effect sizes are substantial: editorially endorsed products see selection probability increases of 14–43% in some categories, while 'sponsored' or promotional tags produce decreases of 8–13%. This maps directly to landing page copy: claims backed by named sources, specific numbers, or third-party validation outperform equivalent claims that are self-asserted.

Choice homogeneity: AI buyers collapse to consensus

Human consumers distribute purchase intent across a range of options. AI buyers, evaluated at scale, show significantly higher consensus: demand collapses onto one or two dominant products, with the long tail of alternatives receiving near-zero selection probability. This is not a problem to eliminate — it is a signal to interpret. When eLLMo's panel of personas reaches strong consensus around a specific friction point, that consensus is a reliable indicator of a real conversion barrier. When the panel diverges, the insight is about segment-specific positioning. Consensus and divergence are both informative.

Model version shifts are distributional, not incremental

When a frontier model releases a new version, the behavioral profile of AI buyers changes in ways that cannot be predicted by extrapolation from the prior version. In documented cases, position preferences inverted (bottom-row bias became top-row bias after a model update), and price elasticity coefficients shifted significantly. For a simulation product, this creates a requirement: every simulation run must be tagged with a specific model version, and version upgrades must be treated as a change in the measurement instrument, not just an improvement to it.

The experiment that changed the question

In 2025, researchers built ACES — the Agentic e-CommercE Simulator — to answer a question the industry had been circling for years: what does a frontier AI model actually do when it sits in the buyer's seat? The setup was controlled and deliberate. Frontier shopping agents were placed inside a mock storefront. Thousands of randomized trials varied product position, price, ratings, review counts, and badge labels. The agents selected products. The researchers measured which variables drove those selections, how strongly, and whether the patterns held across model families.

The finding was not that AI buyers are irrational. It was that they are rational and biased simultaneously — and that the biases are structural enough to quantify, consistent enough to predict, and meaningful enough to build a product around.

Rational in aggregate, biased in the detail

ACES established that economic rationality is no longer the question. Frontier models, when instructed to select the cheapest product, select the cheapest product. Their price elasticity — how selection probability responds to a price change — falls between -1.6 and -2.8, a range comparable to published human consumer benchmarks. Rating elasticity operates in the same correct direction. Agents are not confused about what value means.

What ACES surfaced beneath that rationality was a second layer: systematic preferences that do not follow from value alone, that persist across different presentation formats, and that vary between model families in ways the models themselves cannot explain. These are not bugs in isolated prompts. They are behavioral properties of how language models process information — properties that show up every time an AI agent evaluates an ordered set of options.

That combination — rational preference structure plus consistent structural biases — is precisely what makes AI-buyer simulation tractable. A purely rational system would tell you nothing beyond price and rating sensitivity. A purely random system would tell you nothing at all. The real situation is that AI buyers respond correctly to value signals while also responding systematically to information structure, credibility framing, and presentation order. Each of those response patterns is a dimension that eLLMo maps onto a landing page.

Position bias: the mechanism behind the order effect

The most counterintuitive ACES finding is position bias — and the counterintuitive part is not that it exists, but why. Agents heavily favor products in certain list positions, most prominently top-of-list placements. That pattern reads at first like a visual attention effect: the agent sees the top item first, lingers on it, selects it. ACES ruled that out directly by running the same experiments in pure JSON and plain-text interfaces, with no visual rendering at all. The bias persisted.

The mechanism is in the model's training architecture. Language models are trained on human-generated text in which the first item in a list tends to be the most important, the most recommended, or the default choice. That pattern is encoded as a prior. When the model processes an ordered sequence of products, items in the first positions receive implicit weight from that prior, independent of the content at those positions. The bias is not visual. It is how a model weights an ordered sequence.

Translated to a landing page, this has a direct implication for information hierarchy. The content a buyer encounters first in the document's linear flow — the headline, the hero claim, the first paragraph of body copy — carries disproportionate weight in the evaluation that follows. An objection addressed early is resolved before it compounds. Third-party proof that sits in the second scroll position reaches the buyer before a price that sits in the third. Sequence is not just a design choice. It is a variable in the buyer's reasoning.

Framing and anchoring: training data as behavioral inheritance

ACES confirmed that AI buyers replicate the framing and anchoring effects documented in human behavioral economics research. The same product option presented as a gain draws a different response than the same option presented as a loss. A reference price pulls subsequent valuations toward it. An option positioned as the established default in a set receives elevated selection probability without any change to its attributes.

These effects are not programmed into frontier models as explicit rules. They are inherited from training data: the models have processed enough human-generated text — reviews, comparisons, purchasing guides, consumer forums — that they have absorbed the behavioral patterns humans exhibit when evaluating choices. The effects appear in the right direction in ACES experiments, sometimes attenuated relative to human magnitude, but consistent enough to measure and consistent enough to test against.

For a DTC landing page, framing is operative at every level. A hero claim that leads with what the buyer gains ('recover two hours a week') activates a different evaluation pathway than one that leads with what the buyer currently lacks. An anchor price that establishes a reference before the product price is presented shapes how that price is perceived. A call to action that presents the purchase as the default option in a familiar choice pattern performs differently than one that asks the buyer to initiate from a cold start. eLLMo surfaces these framing signals as a dedicated dimension in every simulation run — not as general copy advice, but as findings tied to the specific language on the specific page.

Credibility signals: what AI buyers treat as earned versus asserted

One of the most operationally useful ACES findings is the asymmetry between credibility types. In product selection experiments, editorial and third-party endorsement signals — badges labeled 'Overall Pick,' press mentions, verified external reviews — raised selection probability by 14–43% across some product categories. Promotional signals — badges labeled 'Sponsored,' brand-asserted superlatives, self-assigned category leadership — reduced selection probability by 8–13%.

The mechanism is consistent with how models encode the value of information. A claim attributed to a third party has a different epistemic status than a claim a seller makes about themselves. Models trained on human text have processed the pattern that editorial sources carry more weight than promotional ones — that a press mention is harder to manufacture than a tagline, that a specific number cited by an external source is harder to dispute than an unsupported 'best.' The AI buyer inherits that discrimination and applies it faster and more consistently than a skeptical human would.

  • Rewarded: named publication mentions, award attributions, specific performance numbers from third-party tests, customer quotes with verifiable detail, analyst endorsements.
  • Penalized: unsubstantiated superlatives ('the leading solution'), self-assigned 'best' claims without named comparison, promotional framing without underlying evidence.

The practical reading for a landing page is this: a claim a buyer can check outperforms an equivalent claim a buyer has to take on faith. That is not a stylistic preference — it is a measurable selection effect.

'The hero claim is clear. But for high-skepticism buyers, there is no external source that corroborates it — and that gap is where the evaluation stalls.'

Choice homogeneity: what consensus means in simulation

Human consumers distribute purchases across product assortments. Given a shelf of ten options with differentiated prices and attributes, they spread. AI agents evaluated at scale do not. ACES documented choice homogeneity as a consistent behavioral property: demand collapses onto one or two modal products, with selection probability near zero for the rest of the assortment. The distribution is winner-take-all in a way that human purchase behavior is not.

For a buyer simulation product, this property carries two distinct implications. The first is a calibration boundary: eLLMo's persona panel should not be expected to reproduce the full distributional spread of real buyer behavior across a population. The long tail of niche objections and minority preferences is underrepresented when AI buyers converge. The second implication is more useful — consensus is a signal. When a calibrated panel of differentiated personas, each conditioned on a distinct buyer profile, converges on the same friction point, that convergence is a reliable indicator of a real barrier. It is not a modeling artifact. When 8 of 10 personas stall at the same element, the element has a structural problem that transcends segment.

The complement is also true. When the panel splits — when high-skepticism personas flag a credibility gap that price-sensitive personas do not register — the finding is about segment-specific positioning. The friction is real but not universal, and the fix is in the messaging sequence for that segment, not in the global structure of the page. Consensus and divergence each carry diagnostic meaning. eLLMo reports both, because the split is as actionable as the convergence. See also the preference heterogeneity research for how persona conditioning produces segment-coherent variation.

Model version as the measurement instrument

ACES produced a finding that sits outside the behavioral findings but has direct implications for any product built on top of frontier models. When a model ships a new version, the behavioral profile of AI buyers changes in ways that are not incremental. Documented cases in the research include position preferences inverting completely between two versions of the same model family, and between preview and official releases of a different family — a shift from bottom-row preference to top-row preference. Price elasticity coefficients moved significantly across versions. These are not minor perturbations in a stable signal. They are what the researchers call demand shocks: the same product assortment, evaluated by the same agent type, produces a completely different purchase distribution after a model update.

The structural implication is that model version is the measurement instrument. A simulation run on one model version and a simulation run on a newer version of the same model are not comparable without a calibration layer that accounts for the behavioral shift. That is not a weakness specific to eLLMo. It is a property of the technology. The operational response is to treat it as a first-class variable rather than a background assumption.

eLLMo pins model version to every simulation run. A version change to the simulation engine is treated as a change to the instrument, not a free upgrade, and clients receive a delta report that separates changes in landing-page performance from changes in the underlying behavioral baseline. The version record is what makes simulation findings reproducible — and reproducibility is what makes them comparable over time.

How eLLMo operationalizes each bias dimension

The ACES findings map directly to the dimensions eLLMo evaluates on every landing page. Position bias becomes an information-order analysis: does third-party proof appear before the price, or after it? Does the hero claim establish the buyer's frame before objections can form, or does the page lead with a category claim that requires the buyer to fill in the relevance themselves? The simulation marks where in the document's linear structure confidence breaks for each persona type.

Framing and anchoring effects become a copy-structure analysis: does the page lead with gain or with absence, and does that framing match the motivational orientation of the personas most likely to convert? Are reference prices present, and do they anchor correctly for the target price point? Credibility signal analysis maps directly from the ACES badge experiment: every claim on the page is evaluated for whether it is attributed to a named source a buyer can verify, or asserted by the brand without external grounding. Choice homogeneity informs how eLLMo reads panel consensus — a finding that appears across 8 or more personas in a calibrated panel is flagged as a structural barrier, not a segment-specific preference. And model version is pinned and reported on every run, so the instrument is never a silent variable in the output.

The result is a simulation report that does not return abstract copy advice. It returns findings like: 'the only external credibility signal on this page appears at scroll position four, below the price reveal — skeptical personas have already discounted the claim before they reach the evidence that would support it.' That is a finding of sequence, not of substance. The page has the evidence. The order is wrong. That is the kind of specific, actionable output that simulation validity research shows correlates with real conversion improvement when acted on.

Methodology note

Built on the research. Designed for decisions.

eLLMo simulation surfaces ranked friction patterns across calibrated buyer personas — specific findings, traceable to buyer segments, actionable on the same day. The methodology is grounded in peer-reviewed research on AI agent behavior and OCEAN psychometrics. The output is a prioritized list of what to fix before your campaign launches — and why it matters for each buyer type.