Research Area 05

AI Agents as Economic Actors: What Live Negotiation Reveals About Simulation

Anthropic's Project Deal experiment didn't set out to validate buyer simulation — but it produced the most direct empirical evidence to date that AI agents authentically encode economic reasoning, not just approximate it. Sixty-nine employees handed their preferences to AI agents, which then negotiated and executed 186 real trades — $4,000 in total value, no human intervention — in natural language across Slack. The findings speak directly to simulation methodology: what makes an AI agent a valid proxy for human economic decision-making, and what breaks that validity silently.


Key takeaways

AI agents execute real economic transactions through natural language reasoning

Project Deal agents identified potential trading matches, proposed prices, negotiated counteroffers, and closed deals across a diverse transaction set — physical goods, experiential trades like dog-sitting, and services — without pre-defined negotiation protocols. 186 deals closed. $4,000 in total value. No human in the loop at any point. This is not a controlled laboratory finding: it is evidence that AI agents can authentically represent human economic interests in open-ended, multi-party settings. For buyer simulation, this matters because it validates the foundational premise: an AI agent reasoning in natural language about a landing page is doing something economically real, not producing a plausible-looking facsimile of evaluation.

Simulation instruments fail silently — and users don't notice

Project Deal's most consequential finding is not the performance gap between Opus and Haiku agents — it is that participants using the weaker model rated deal fairness identically to Opus users (approximately 4 on a 7-point scale), despite completing objectively fewer deals and extracting less value per sale. The output looked right. The experience felt fair. The gap was invisible from inside. For DTC teams using AI simulation to evaluate landing pages, this is the critical calibration question: how would you know if your simulation instrument is systematically missing signal? A weak-model evaluation produces plausible findings, surfaces real-sounding friction, and returns a structured report — and it may have missed the friction patterns that matter most. The failure mode does not announce itself.

The reasoning trace, not the verdict, is where buyer psychology lives

Project Deal agents negotiated in natural language — proposing, countering, reasoning about value. Prompting agents to negotiate aggressively had no statistically significant effect on outcomes. The model's embedded sense of value — what it understood about the transaction, the items, the context — determined behavior far more than the instruction layer did. This maps directly to simulation methodology: the actionable output of an AI buyer evaluation is not the pass/fail verdict or conversion probability. It is the reasoning trace — why the agent hesitated, what it weighted, what it misread. That trace is more useful for landing page copy decisions than any score. A finding like 'the agent paused because the price anchor appeared before the value justification, which it interpreted as an attempt to pre-empt objections' tells a copywriter exactly what to reorder. A conversion rate does not.

A live test of agents acting as economic principals

Most evidence about AI agent behavior comes from synthetic benchmarks — tasks designed to be measurable, not conditions designed to be real. Anthropic's Project Deal ran the experiment differently. Sixty-nine employees each handed their economic preferences to an AI agent, then stepped back entirely. Those agents negotiated and executed real trades with each other across a Slack-based platform, no human in the loop at any point, closing 186 deals worth roughly $4,000 in total value. The transaction set was deliberately heterogeneous: physical goods, professional services, experiential trades — dog-sitting among them — chosen specifically because the value of each item was genuinely subjective and context-dependent. The agents had to reason about worth, not retrieve a price from a table.

The result is not a laboratory finding. It is evidence that an AI agent reasoning in natural language about a transaction is doing something economically real — representing interests, weighing tradeoffs, adapting to counteroffers, and closing at a price. For anyone building buyer simulation on the premise that an AI agent can stand in for a human economic decision-maker, this is the external validation the field needed. The agent is not producing a plausible imitation of economic reasoning. It is executing it.

The silent failure: the finding that changes how to think about calibration

The headline result — agents close real deals, at real value, across open-ended categories — is useful. The result that should change product and methodology decisions is quieter, and more consequential.

Project Deal ran agents on two different models: a stronger one and a weaker one. The performance gap was real and measurable. Agents on the stronger model completed more deals and extracted more value per transaction. Agents on the weaker model completed fewer deals and left value on the table. By any external measure, the two populations produced different economic outcomes.

Participants rated those outcomes identically.

Both groups rated the fairness of their deals at approximately 4 on a 7-point scale. The weaker-model group did not report worse experiences. They did not flag confusion or dissatisfaction. Their subjective read of how the negotiation went matched the subjective read of the stronger-model group, point for point — while the objective quality of their outcomes diverged. The gap was invisible from inside the experience.

Plausible output and valid output look identical from the inside. Only calibration tells them apart.

This is the silent failure: a simulation instrument that returns structured results, coherent reasoning, and a confident report — while systematically missing the signal that matters. The failure does not surface as an error message or an incoherent output. It surfaces as a finding that looks right, reads right, and is wrong in ways the team running the evaluation cannot detect. That failure mode is not hypothetical. Project Deal documented it in a controlled experiment with real stakes and real participants.

For any team using AI simulation to evaluate landing pages, this is the calibration question made concrete: if your evaluation instrument were running on a weaker model, how would you know? The output would still arrive. The friction findings would still sound plausible. The report would still be structured. Nothing in the surface of the result announces the gap.

What drove behavior: embedded value, not the instruction layer

Project Deal also tested a natural assumption about how to improve agent performance: instruct the agent to negotiate harder. If you want better outcomes, tell the agent to be more aggressive. The experiment ran that condition directly — prompting agents to negotiate with assertive, value-capturing intent — and measured the difference.

There was no statistically significant difference.

Outcome quality did not move when the instruction layer changed. What moved outcomes was the model itself — specifically, what the model understood about the transaction, the items, the context, and what counted as a good deal. The model's embedded sense of value, built from training rather than from the system prompt, determined how agents actually behaved.

The instruction layer is real. It shapes behavior in identifiable ways — see human heterogeneity in agentic markets for the evidence on how much prompt authorship drives outcome variance. But the embedded knowledge that makes an agent economically competent is not something a prompt installs. It is something the model either has or does not have. Telling an agent to negotiate aggressively does not give it better judgment about what constitutes a good outcome. The judgment either lives in the model, encoded from training, or it does not.

This has a direct implication for simulation methodology. If you want an AI buyer to surface the friction that a real buyer would find, the instrument needs a model whose embedded understanding of value, credibility, and buyer psychology is calibrated to the task. Prompting a weaker model harder does not close that gap. It produces a more assertive version of the same weak evaluation.

The reasoning trace is the deliverable

Project Deal's agents negotiated in natural language. They did not score items on a rubric or return a probability estimate. They proposed, countered, justified, and responded to counteroffers — in sentences, following the same conversational logic a human would use to work through a deal.

That structure is the methodology. The valuable output of an AI buyer evaluation is not a verdict, a pass/fail, or a conversion probability. It is the reasoning trace: the record of what the agent weighed, where it hesitated, what it found credible, and what it found unconvincing. A reasoning trace says something a score cannot. It says why.

A finding in that form reads like a buyer talking back to the page: 'The price anchor appeared before the value justification, and I read that as an attempt to pre-empt the objection rather than answer it — so I discounted both.' That finding tells a copywriter exactly what to reorder. It identifies the sequence that created the problem, the inference the buyer drew from it, and the consequence for trust. A conversion rate is a measurement of something that already happened. A reasoning trace is an explanation of why it happened and what to change.

When reasoning traces are aggregated across a panel of calibrated personas, they produce a friction map: which hesitations appeared in every segment, which concentrated in specific buyer types, and which were isolated. The distribution across the panel is itself information. A finding that appears in 9 of 10 personas identifies a universal barrier. A finding concentrated in high-skepticism or high-conscientiousness segments identifies a positioning opportunity for that buyer type.

How eLLMo operationalizes these findings

Project Deal validates three specific design choices that eLLMo is built around.

The first is that the reasoning trace is the primary output. eLLMo does not return a conversion probability or a page score. It returns the reasoning behind each persona's response to the page — what it weighted, where confidence built, where it broke. That trace is what a product, copy, or design team can act on. The panel's aggregated traces produce a ranked friction map where every finding ties a buyer segment, a page element, and the logic behind the response.

The second is that the model is pinned as a controlled instrument. Because the Project Deal finding shows that weaker models fail silently — returning plausible output that misses real signal — eLLMo treats model version as a first-class provenance dimension. Every simulation run is tagged to a specific model version. When the underlying model changes, the change is logged and surfaced as a shift in the instrument, not absorbed invisibly into results that look continuous. A finding reflects buyer psychology, not an undisclosed change in what did the evaluating.

The third is designing against the silent failure mode. Silent failure — the result that looks right while missing the signal — is the most expensive failure in simulation, because it does not trigger correction. eLLMo's calibration methodology addresses this directly: cross-run consistency checks, persona-level reasoning audits, and version-to-version delta reporting are all built to surface the gap between plausible output and valid output before it compounds into a wrong decision. The simulation validity research area describes the calibration methodology in full.

The commercial implication: the instrument choice is a product decision

Project Deal ran when agent-mediated commerce was still in early deployment. The commercial trajectory since then has been directional and fast. AI shopping agents that browse, compare, and transact on a consumer's behalf are moving from demonstration to default — which means the buyer reading your landing page is increasingly an agent, not a person.

That shift makes the instrument choice a commercial decision, not a technical one. A brand that tests its page against a calibrated AI buyer — running on a model whose embedded understanding of value and credibility is strong enough to surface what a real buyer finds — is reading the conversion surface that matters. A brand that tests against a weaker instrument, or does not test at all, is operating on the assumption that their page works for buyers they have never actually interrogated.

Project Deal's silent failure finding is the sharpest version of that risk. The participants who ran weaker agents thought their deals were fair. They were not poorly informed or careless. They had no way to see the gap from inside the result. The brands that treat simulation instrument quality as a decision worth making — and calibration as a discipline worth maintaining — are building the one thing that defends against the most expensive failure mode: the confident answer that was wrong all along.

Methodology note

Built on the research. Designed for decisions.

eLLMo simulation surfaces ranked friction patterns across calibrated buyer personas — specific findings, traceable to buyer segments, actionable on the same day. The methodology is grounded in peer-reviewed research on AI agent behavior and OCEAN psychometrics. The output is a prioritized list of what to fix before your campaign launches — and why it matters for each buyer type.