Quantifying Prior Dominance in RAG Systems
Quick Answer
The study introduces the Normalized Context Utilization (NCU) metric to evaluate Retrieval-Augmented Generation (RAG) systems, revealing that Small Language Models (SLMs) outperform larger models in factual extraction.
Quick Take
The findings indicate that traditional scaling laws yield diminishing returns, with a commercial API frequently failing against adversarial evidence due to systemic confidence collapse.
Key Points
- Normalized Context Utilization (NCU) quantifies contextual information gain in systems.
- Small Language Models (SLMs) outperform larger models in strict factual extraction tasks.
- Commercial API showed systemic confidence collapse when contradicted by external evidence.
- Prior Dominance correlates with model scale and proprietary alignments.
- Traditional scaling laws exhibit extreme diminishing returns in large models.
DeepSignal Analysis
What happened
The study introduces the Normalized Context Utilization (NCU) metric for evaluating Retrieval-Augmented Generation (RAG) systems. It finds that Small Language Models (SLMs) can outperform larger models in factual extraction, challenging traditional scaling laws. Additionally, a commercial API often fails against adversarial evidence due to systemic confidence collapse.
Key evidence
- The Normalized Context Utilization (NCU) metric quantifies contextual information gain using continuous token log-probabilities across various conditions.
- The research shows that Small Language Models (SLMs) can match or outperform larger models, indicating diminishing returns from traditional scaling laws.
- The evaluated commercial API frequently contradicted explicit external evidence in nearly half of adversarial conflicts, leading to systemic confidence collapse.
Why it matters
These findings suggest that smaller models may be more effective for specific tasks like factual extraction, which could influence future model development and deployment strategies. The performance of the commercial API raises concerns about its reliability in adversarial contexts, highlighting the need for improved evaluation metrics in RAG systems.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristics that suffer from ''epistemic blindness'' - failing to distinguish genuine contextual information extraction from parametric memory recall. To address this, we introduce the Normalized Context Utilization (NCU) metric, leveraging continuous token log-probabilities across zero-shot, oracle, and adversarial conditions to strictly quantify contextual information gain. Evaluating architectures ranging from 1.5B to 72B parameters alongside a proprietary commercial API reveals that for strict factual extraction (without Chain-of-Thought reasoning), traditional scaling laws exhibit extreme diminishing returns: highly efficient Small Language Models (SLMs) match or outperform high-capacity architectures. Furthermore, we demonstrate that ``Prior Dominance'' correlates with model scale and proprietary alignments. The evaluated commercial API not only overrode explicit external evidence in nearly half of adversarial conflicts, but also frequently suffered from systemic confidence collapse (Negative Transfer) when its parametric priors were contradicted. Our findings highlight the structural epistemic advantage and superior contextual adherence of SLMs in strict extraction workflows.
| Comments: | 15 pages, Preprint |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.23695 [cs.CL] |
| (or arXiv:2606.23695v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.23695 arXiv-issued DOI via DataCite |
Submission history
From: Barak Or [view email]
[v1]
Wed, 29 Apr 2026 15:38:24 UTC (89 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.