Which RAG Paradigm Wins at Scale? A Scaling Study of Retrieval-Augmented Generation Paradigms
Quick Answer
This paper shows that In a controlled scaling study of retrieval-augmented generation (RAG) paradigms, BM25 outperformed others by defining the low-cost end of the Pareto frontier and leading accuracy from mid-scale onward.
Quick Take
The File-System Agent initially matched BM25 but fell significantly behind at larger scales, while graph-based struggled with construction efficiency and accuracy. Overall, BM25 remains the most effective method for large-scale applications.
Key Points
- BM25 defines the low-cost end of the Pareto frontier across all measured tiers.
- File-System Agent uses 39 times more query tokens at the bedrock than BM25.
- Agent+BM25 scores 69.4 at full scale, outperforming raw-file agency and native BM25.
- Graph-based RAG struggles with construction, using excessive tokens with low accuracy.
- Study spans 28 tiers from 1,000 to 512,000 documents, maintaining fixed questions.
Paper Resources
Source Excerpt
(RAG) methods range from lexical and dense retrieval to graph-based indexing and agentic search. They are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled corpus-scaling study of these four paradigms. A ladder of 28 strictly nested tiers grows from roughly 1,000 to 512,000 documents while questions and a fixed bedrock of relevant and adversarial documents remain un
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.