When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
Quick Answer
This study evaluates LLM-based urban simulators like AgentSociety and CitySim, revealing a significant gap between narrative plausibility and real-world mobility realism.
Quick Take
Using datasets from Greater Paris and Shanghai, the analysis shows these models struggle with core spatial and temporal constraints, necessitating rigorous empirical validation and improved initialization methods for realistic urban simulations.
Key Points
- Introduces a validation framework for evaluating mobility realism in -based simulators.
- Finds substantial gaps in trip-length distributions and origin-destination flows.
- Highlights instability in realistic mobility diversity across prompting configurations.
- Provides scalable infrastructure for map generation and mobility-metric computation.
- Emphasizes the need for empirical validation in urban simulation systems.
Paper Resources
Source Excerpt
arXiv:2606. 13835v1 Announce Type: new Abstract: -based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic human mobility patterns or merely generate plausible mobility narratives. We introduce a validation framework for evaluating the mobility of generative agents of LLM-based urban simulators against real-world mobility data.
For this, we use mobility laws, temporal rhythms, network motifs, semantic activity transitions, and behavioral mobility profiles. Using datasets from the Greater Paris region and Shanghai, we evaluate AgentSociety and CitySim across multiple dimensions of mobility realism. …
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.