OdysSim: Building Foundation Models for Human Behavior Simulation
Quick Answer
OdysSim introduces a novel 8B OSim model, outperforming existing models on 8 out of 23 tasks, particularly in conversational and social simulations.
Quick Take
The study highlights the need to rethink training paradigms to bridge the Sim2Real gap and improve human-like interaction quality.
Key Points
- OdysSim corpus includes 21.4M interactions and 10B tokens for training.
- SOUL taxonomy unifies 62 datasets and 23 benchmark tasks into one framework.
- OSim model achieves 93.2 alignment with real users on reaction tasks.
- Post-training reward-hacking patterns are mitigated using specialized detectors.
- All research artifacts are released to support future investigations.
Paper Resources
Source Excerpt
arXiv:2606. 14199v1 Announce Type: new Abstract: are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i. e. , models trained to simulate human behavior at scale.
We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. …
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.