Academic foundation

The science behind
eLLMo simulation.

The research below establishes how AI buyers actually read a page — where they hesitate, what they trust, and how their behavior shifts with the model. eLLMo turns that into a measurement instrument: calibrated buyer personas that surface ranked friction before you spend. This is the foundation, without the footnotes.

Research topics

Research Area 01

AI Buyer Behavior and Structural Bias

A substantial body of controlled research has tested how frontier AI models behave when placed in the role of autonomous buyers — making product selections, comparing options, and evaluating landing page content. The consistent finding is that AI buyers are economically rational in aggregate but exhibit systematic, non-human biases in how they process information.

Key takeaways
  • Position effects are structural, not visual
  • Framing and anchoring effects are measurable and exploitable
  • Credibility signals vs. self-assertion: a persistent divide
  • Choice homogeneity: AI buyers collapse to consensus
  • Model version shifts are distributional, not incremental
Research Area 02

Preference Heterogeneity and Persona Validity

The foundational question for any AI simulation product is whether AI personas can reliably represent how different types of human buyers respond to marketing stimuli. Research answers this clearly: AI models reliably surface meaningful variation across simulated buyer types — and eLLMo is built around that capability. Understanding which buyer segments respond differently, and why, is the insight that determines what to fix and what to leave alone.

Key takeaways
  • Persona conditioning produces demographically coherent responses
  • Conjoint-style preference elicitation yields economically meaningful outputs
  • Why persona conditioning produces better signals than direct elicitation
  • Persona-differentiated variation is durable
  • Chain-of-thought prompting partially corrects absolute calibration
  • Language structure affects AI preference signals
Research Area 03

The OCEAN Framework in Buyer Simulation

The Big Five personality model — Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism — is the most rigorously validated personality framework in psychology, with decades of research linking it to consumer behavior, risk tolerance, information processing, and purchase decision patterns. eLLMo applies OCEAN calibration to every synthetic buyer persona, not as a label but as a behavioral parameter set that shapes how each persona processes landing page information.

Key takeaways
  • Openness and novelty processing
  • Conscientiousness and information completeness
  • Neuroticism and risk signal detection
  • Empirical purchase behavior mappings by OCEAN dimension
  • Cross-persona pattern analysis as the primary signal
Research Area 04

Simulation Validity and Calibration Methodology

AI buyer simulation is more useful than a premature conversion-rate prediction. It shows why different buyers hesitate, what evidence they need, and which page changes are most likely to remove friction before paid traffic is spent — each finding tied to a specific buyer segment and a specific element on the page.

Key takeaways
  • Ranked friction patterns, traceable to buyer segments
  • Hypothesis generation, then validation
  • Credibility-vs-self-promotion as a simulation dimension
  • The proprietary calibration gap
Research Area 05

AI Agents as Economic Actors: What Live Negotiation Reveals About Simulation

Anthropic's Project Deal experiment didn't set out to validate buyer simulation — but it produced the most direct empirical evidence to date that AI agents authentically encode economic reasoning, not just approximate it. Sixty-nine employees handed their preferences to AI agents, which then negotiated and executed 186 real trades — $4,000 in total value, no human intervention — in natural language across Slack. The findings speak directly to simulation methodology: what makes an AI agent a valid proxy for human economic decision-making, and what breaks that validity silently.

Key takeaways
  • AI agents execute real economic transactions through natural language reasoning
  • Simulation instruments fail silently — and users don't notice
  • The reasoning trace, not the verdict, is where buyer psychology lives
Research Area 06

Human Heterogeneity in Agentic Markets

Delegating a decision to an AI agent does not strip out human difference — it transmits it. In a controlled study of AI-mediated negotiations, identical models pursuing identical objectives produced widely dispersed outcomes, and most of that dispersion traced to the human who wrote the agent's instructions rather than to the model itself. The prompt carried the signal. That result is the empirical foundation for persona conditioning: how a buyer is specified is what determines how it behaves.

Key takeaways
  • Delegation preserves human difference — and widens it
  • The prompt is a transmission mechanism for identity
  • Machine fluency is a new, measurable form of human capital
  • Social norms erode under delegation
  • Specification hazard replaces information asymmetry
Research Area 07

Agent-to-Agent Commerce and Model-Dependent Outcomes

The buyer on the other side of your page is increasingly an agent — and which agent it is changes the outcome. A benchmark of agent-to-agent negotiation in consumer markets, where both the shopper and the merchant delegate to AI agents that negotiate price and close the deal without a human in the loop, found that the model behind the buyer is a first-order determinant of who captures value. Automated commerce is not a level playing field. It is an imbalanced game decided in part by model choice.

Key takeaways
  • Automated deal-making is imbalanced by model
  • Weaker agents fail expensively
  • Behavioral anomalies become real losses
  • Model choice is now a demand-side variable
  • A simulation is only as valid as its benchmarked agent
Research Area 08

The Productivity Compression and the Pre-Spend Advantage

AI compresses the time it takes to do skilled work, and the harder the task, the larger the compression. Analyzing 100,000 real conversations with an AI assistant, Anthropic estimated that tasks which would take about 90 minutes unaided were completed roughly 80% faster with AI assistance, with the steepest gains on the most cognitively demanding work. Pre-spend buyer simulation applies that same compression to the slowest, most expensive part of conversion work: finding out what breaks before you pay for the traffic that finds out for you.

Key takeaways
  • AI collapses task time on skilled work
  • The most complex work compresses the most
  • Adoption could double labor-productivity growth
  • The conversion-research cycle is a prime compression target
  • The return is decision latency, not just labor saved
Research Area 09

The AI-Shaped Buyer: Evidence from 81,000 Interviews

The buyer landing on your page now arrives with a formed relationship to AI, and the largest qualitative study of that relationship to date maps what they bring with them. Over one week, Anthropic interviewed roughly 81,000 people across 159 countries and 70 languages using an AI interviewer, then classified the responses at scale. The result is a population-level picture of hope, fear, and trust around AI — and a working demonstration that structured AI inquiry at scale produces decision-grade signal.

Key takeaways
  • The buyer's relationship to AI is defined by paired tensions
  • Productivity is felt, but uneven
  • Trust and anxiety travel with the buyer
  • AI inquiry at population scale is now a validated method
  • The panel mirrors the population
Research Area 10

Staying in Character: Whether a Simulated Buyer Holds Across a Run

Every other area here asks whether a simulated buyer produces useful signal. This one asks a question underneath that: did the buyer stay the same person from the first turn to the last? Researchers at UC Berkeley, the University of Washington, and Google DeepMind built three automatic checks for that, tested them against thirty human raters, and found that models hold a conversation together far better than they hold a character.

Key takeaways
  • Reading smoothly and staying in character are different measurements
  • The drift does not look like a mistake
  • Length is the stress test
  • The automatic checks were steadier than the people
  • Consistency is trainable, and the gains are not small
Methodology note

Built on the research. Designed for decisions.

eLLMo simulation surfaces ranked friction patterns across calibrated buyer personas — specific findings, traceable to buyer segments, actionable on the same day. The methodology is grounded in peer-reviewed research on AI agent behavior and OCEAN psychometrics. The output is a prioritized list of what to fix before your campaign launches — and why it matters for each buyer type.