
Google Research's Gemini-SQL2 tops text-to-SQL benchmarks by a wide margin
Quick Answer
Google Research's Gemini-SQL2, based on Gemini 3.1 Pro, achieves 80.04% accuracy on the BIRD benchmark, outperforming competitors like OpenAI and Anthropic.
Quick Take
This advancement could enhance natural language processing capabilities across Google's data services.
Key Points
- Gemini-SQL2 converts natural language into executable SQL queries.
- Achieved 80.04% accuracy on the BIRD benchmark.
- Significantly outperforms OpenAI and Anthropic in text-to-SQL tasks.
- Potential to enhance natural language features in Google's data services.
- Developed on the advanced Gemini 3.1 Pro architecture.
📖 Reader Mode
~1 min readGoogle Research unveiled Gemini-SQL2, a new text-to-SQL system built on Gemini 3.1 Pro. It translates natural language into executable SQL database queries. On the BIRD benchmark, which measures how accurately these translations work, Gemini-SQL2 hits an execution accuracy of 80.04 percent, putting it in first place, according to Google. OpenAI's GPT-5.5-xhigh scores about 72.8 percent, and Anthropic's Claude Opus 4.6 lands around 70.9 percent. Models from Databricks, AWS, Tencent, and Alibaba all trail well behind.

Google Research points out that turning natural language into correct SQL is especially hard because data is often layered and queries need to account for complex business logic. The generated SQL queries both look correct and execute successfully, the company says.
Better SQL understanding could improve natural language features across Google's data services more broadly, according to Google. The research team hasn't said anything about a public release of the model, and there's no paper yet either.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

