ARCS: Towards Precise Text-to-SQL via Structured Disambiguation
Quick Answer
The ARCS benchmark introduces structured disambiguation for text-to-SQL systems, addressing user question ambiguities that lead to errors.
Quick Take
Experimental results show gpt-6-sol achieves only 51% execution accuracy, while no open-source model surpasses 27%. This highlights the challenges in real-world SQL deployments.
Key Points
- ARCS is the first text-to-SQL benchmark with real-world ambiguities.
- Structured disambiguation resolves ambiguities through constrained interactions.
- gpt-6-sol achieves only 51% end-to-end execution accuracy.
- No open-source model exceeds 27% accuracy in text-to-SQL tasks.
- Ambiguities in user questions often lead to significant errors.
Paper Resources
📖 Reader Mode
~2 min readAbstract:As text-to-SQL systems move beyond demonstrations toward real-world deployment, ambiguity in user questions becomes a primary source of errors. Such ambiguities are often subtle, domain- or data-specific, and can silently cause system outputs to deviate from the user's true intent. Ambiguity is traditionally addressed through conversational clarification, which is often inefficient, cognitively demanding, and poorly aligned with real-world user workflows. We propose structured disambiguation, a new paradigm in which ambiguity is resolved through explicit, constrained interactions rather than free-form dialogue. We construct ARCS (Ambiguity Resolution Corpus for SQL), the first text-to-SQL benchmark featuring naturally occurring, unconstrained ambiguities over real-world databases, with complete annotations of all valid ambiguity points, interpretations, and SQL queries. Experimental results show that text-to-SQL remains challenging in the presence of ambiguity: gpt-6-sol achieves only 51% end-to-end execution accuracy, and no open-source model exceeds 27%.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09396 [cs.CL] |
| (or arXiv:2610.09396v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09396 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yanlin Feng [view email]
[v1]
Wed, 7 Oct 2026 03:54:02 UTC (1,398 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.