RAS: Reflection-Augmented Scaling with In-Context Learning for Executable Cypher Query Generation
Quick Answer
The study introduces Reflection-Augmented Scaling (RAS) for generating executable Cypher queries, achieving a 41-50% reduction in Query Execution Error Rate compared to Independent Scaling (IS) at n=5.
Quick Take
This method leverages prior execution feedback through in-context learning, demonstrating that execution errors can provide actionable insights rather than being discarded.
Key Points
- RAS outperforms IS, reducing Query Execution Error Rate by 41-50%.
- Independent Scaling shows a lower error reduction of 32-38%.
- The study utilizes three Neo4j datasets and five code-specialized models.
- Execution errors are treated as actionable feedback for improved query generation.
- In-context learning enhances the efficiency of query code generation.
Paper Resources
Article Excerpt
From source RSS / original summaryarXiv:2605. 22937v1 Announce Type: new Abstract: can reduce errors in structured query generation, but methods to allocate the compute for query code generation remains underexplored. We study Text2Cypher, where language models generate Cypher queries that execute against property graph databases. Non-executable queries constitute a distinct syntactic failure separate from semantic inaccuracy: a syntax error triggers a system-generated error message from the database.
These error messages are typically discarded at inference time rather than leveraged through in-context learning (ICL). We compare two inference methods: Independent Scaling (IS), which performs memoryless resampling, and Reflection-Augmented Scaling (RAS), which conditions each new attempt on prior execution feedback via ICL. Across three Neo4j datasets and five code-specialized language models, RAS reduces the Query Execution Error Rate by 41--50% at n{=}5, outperforming IS at 32--38%.
Execution errors are not merely failures to discard but actionable feedback, and structuring inference-time compute around them is a more efficient path to executability than scaling independent samples.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.