← Back to Digest
Language & NLPApr 9, 2026

Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

A new benchmark finds that today's best LLMs systematically simulate an idealized average person, not real humans with quirks and bad habits.

6.1
Hunch Score
6.3
Academic
5.0
Commercial
5.0
Cultural
HorizonMid (2-5y)
Evidencemedium
Was this useful?

The Thesis

OmniBehavior is the first benchmark for testing AI "user simulators" — models trained to imitate how real people behave across apps, services, and decisions — built entirely from real-world behavioral data rather than synthetic or single-scenario logs. The key finding is structural: large language models (LLMs) don't just make prediction errors, they are biased toward simulating a fictional "average positive person" who is more active, more consistent, and more optimistic than real users. This bias matters because user simulators are increasingly used to test recommendation systems, train customer-service agents, and run synthetic A/B tests — places where over-optimistic fake users would produce dangerously misleading results. The paper also shows that simply giving models longer context windows doesn't fix the problem; performance plateaus well before human-level fidelity. The catch is that this is a benchmark paper, not a fix — it diagnoses the disease without prescribing a cure.

Catalyst

LLM context windows have recently expanded from ~4K to over 1M tokens, prompting serious commercial interest in running long-horizon simulations of user behavior. At the same time, synthetic data pipelines using LLMs as user stand-ins are being adopted in production recommendation and ad-tech systems — raising the stakes for understanding whether those stand-ins are actually representative. OmniBehavior arrives as that adoption curve steepens, before the field has established rigorous standards.

What's New

Prior user-simulation benchmarks — such as those used in dialogue research or app-agent evaluation — tested models in isolated scenarios: a single shopping session, a single conversation, one app at a time. Those setups made it easy to appear competent but masked how behavior drifts and evolves across contexts and over time. OmniBehavior stitches together cross-scenario, long-horizon traces from real users, revealing a category of failure — persona homogenization and loss of long-tail behavior — that earlier narrow benchmarks structurally could not detect.

The Counter

The paper identifies a real and interesting bias, but it does not demonstrate that this bias actually causes downstream harm in any deployed system — it shows structural differences between LLM outputs and real-user logs, not that those differences break real products. The 'positive average person' framing is evocative, but the magnitude of the performance gap is not clearly tied to a practical failure threshold: how much persona homogenization is tolerable for a given application? The benchmark is also constructed from a specific, undisclosed set of real-world platforms, raising questions about whether the behavioral patterns it captures generalize beyond those contexts. Finally, the paper's conclusion — that models plateau as context grows — may reflect prompt engineering limitations or evaluation design choices rather than a fundamental ceiling in LLM capability. Practitioners building synthetic-data pipelines have workarounds, such as persona injection and diversity-enforced sampling, that the paper does not evaluate against.

Longs

  • META — largest operator of recommendation systems that rely on user-behavior modeling
  • TTD (The Trade Desk) — programmatic ad targeting depends on behavioral prediction accuracy
  • MTCH (Match Group) — matchmaking algorithms are sensitive to simulated vs. real preference distributions
  • BOTZ (Global X Robotics & AI ETF) — broad exposure to applied AI infrastructure
  • Palantir (PLTR) — synthetic data and simulation pipelines for enterprise and government clients

Shorts

  • Companies selling LLM-based synthetic user-data products — the paper's findings suggest their outputs carry systematic bias that undermines validity
  • Recommendation-system vendors who use LLM simulators for offline evaluation — flawed simulators produce optimistic offline metrics that don't transfer to live systems
  • Academic groups whose prior isolated-scenario benchmarks are shown to suffer from 'tunnel vision' per the paper's framing

Enablers (Picks & Shovels)

  • Real-world behavioral logging infrastructure (clickstream, session data) from platforms willing to share anonymized traces
  • Long-context LLM APIs (Gemini 1.5, Claude 3, GPT-4o) used as the evaluation substrate
  • Hugging Face evaluation harnesses for standardized LLM benchmarking
  • Differential privacy tooling needed to safely release real behavioral datasets

Private Watchlist

  • Replicant — conversational AI agent company whose simulation fidelity is directly implicated
  • Etched — inference chip startup whose value depends on LLM deployment at scale
  • Gretel.ai — synthetic data generation company whose products face the exact biases described here
  • Imbue — AI reasoning company focused on agent behavior fidelity

Resources

The Paper

The emergence of Large Language Models (LLMs) has illuminated the potential for a general-purpose user simulator. However, existing benchmarks remain constrained to isolated scenarios, narrow action spaces, or synthetic data, failing to capture the holistic nature of authentic human behavior. To bridge this gap, we introduce OmniBehavior, the first user simulation benchmark constructed entirely from real-world data, integrating long-horizon, cross-scenario, and heterogeneous behavioral patterns into a unified framework. Based on this benchmark, we first provide empirical evidence that previous datasets with isolated scenarios suffer from tunnel vision, whereas real-world decision-making relies on long-term, cross-scenario causal chains. Extensive evaluations of state-of-the-art LLMs reveal that current models struggle to accurately simulate these complex behaviors, with performance plateauing even as context windows expand. Crucially, a systematic comparison between simulated and authentic behaviors uncovers a fundamental structural bias: LLMs tend to converge toward a positive average person, exhibiting hyper-activity, persona homogenization, and a Utopian bias. This results in the loss of individual differences and long-tail behaviors, highlighting critical directions for future high-fidelity simulation research.

Synthesized 4/26/2026, 11:22:32 PM · claude-sonnet-4-6