Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
Quick Answer
This paper shows that In a controlled scaling study of retrieval-augmented generation (RAG) paradigms, BM25 outperformed others by defining the low-cost end of the Pareto frontier and leading accuracy from mid-scale onward.
Quick Take
The File-System Agent initially matched BM25 but fell significantly behind at larger scales, while graph-based struggled with construction efficiency and accuracy. Overall, BM25 remains the most effective method for large-scale applications.
Key Points
- BM25 defines the low-cost end of the Pareto frontier across all measured tiers.
- File-System Agent uses 39 times more query tokens at the bedrock than BM25.
- Agent+BM25 scores 69.4 at full scale, outperforming raw-file agency and native BM25.
- Graph-based RAG struggles with construction, using excessive tokens with low accuracy.
- Study spans 28 tiers from 1,000 to 512,000 documents, maintaining fixed questions.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Retrieval-augmented generation (RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled corpus-scaling study of these four paradigms. A ladder of 28 strictly nested tiers grows from roughly 1,000 to 512,000 documents while questions and a fixed bedrock of relevant and adversarial documents remain unchanged. Under one reader and judging protocol, we measure official accuracy, construction and query tokens, and latency. Our experimental results show that BM25 scales best in this controlled setting: it defines the low-cost end of the Pareto frontier at every measured tier and leads accuracy from mid-scale onward, without LLM-based construction. The File-System Agent matches or slightly exceeds BM25 at the smallest tiers but uses 39 times more query tokens per answer at the bedrock and falls nearly 20 points behind at full scale. A matched retrieval swap reverses this failure: Agent+BM25 scores 69.4 at full scale, versus 36.9 for raw-file agency and 54.8 for native BM25 on the same 150 questions. Graph-based RAG hits a construction wall: its heaviest builders use up to 24.6 generative LLM tokens per indexed corpus token yet stop within the first 2% of the full corpus, while scalable variants remain less accurate than BM25 at shared tiers.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.26497 [cs.CL] |
| (or arXiv:2607.26497v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26497 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pengyu Wang [view email]
[v1]
Wed, 29 Jul 2026 05:46:11 UTC (984 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.