ToolRACER: A Robust Agentic Conversation Emulation Resource for Agent Training and Evaluation
Quick Answer
ToolRACER introduces a synthetic data generation pipeline for training task-oriented conversational agents, creating ToolRACERBench with 5.6K conversation trajectories, 66% of which include failure-prone scenarios.
Quick Take
Models trained on this benchmark show improved accuracy on function-calling benchmarks like τ²-bench and ACEBench, enhancing agent capabilities significantly.
Key Points
- ToolRACER generates multi-turn interactions for robust agent training.
- ToolRACERBench includes 5.6K validated conversation trajectories.
- 66% of conversations in the benchmark feature failure-prone scenarios.
- Models trained on ToolRACERBench outperform others in agentic accuracy.
- Significant improvements noted when mixed with in-domain datasets.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Task-oriented conversational agents remain fragile under real world conversation scenarios as they rarely follow a predictable script, especially when users exhibit non-cooperative behavior. Existing function-calling benchmarks often emphasize successful, cooperative interactions and underrepresent adversarial conversation trajectories, thereby limiting the training resources available for developing robust agents. We present ToolRACER, a synthetic data generation pipeline that coordinates user, assistant and tool emulation models to generate and validated multi-turn interactions between a user and an agent. Using \sysn, we construct ToolRACERBench a robust multi-turn conversation benchmark spanning six domains, ranging over 55 varied personas, generating a validated corpus of 5.6K conversation trajectories, with approximately 66\% of conversations containing failure-prone conversation scenarios. We inject adversarial behaviors, producing validated conversational interaction trajectories that capture realistic, robust scenarios. We evaluate models trained on ToolRACERBench against internal benchmarks, as well as on function calling benchmarks such as $\tau^2$-bench, BFCLv3 and ACEBench to evaluate agentic accuracy and robustness. Models trained on ToolRACERBench improve end to end agentic accuracy across $\tau^2$-bench and ACEBench, demonstrating significant gains when mixed with in-domain dataset in small language models for agent capability tasks.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.09163 [cs.CL] |
| (or arXiv:2610.09163v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09163 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Arkajyoti Chakraborty [view email]
[v1]
Tue, 6 Oct 2026 22:06:17 UTC (1,424 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.