Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
Quick Answer
This study introduces a source-aware reranking method for Retrieval-Augmented Generation (RAG) that incorporates source reliability priors, improving Precision@5 from 0.48 to 0.72 on a health-domain corpus.
Quick Take
The approach reweights retrieval scores based on document source credibility, effectively reducing adversarial document retrieval risks. Experiments were conducted on the Rosie high-performance computing cluster.
Key Points
- Source-aware reranking improves Precision@5 from 0.48 to 0.72.
- Retrieval scores are reweighted using domain-informed source reliability priors.
- Method reduces average adversarial document retrieval risks.
- Experiments conducted on the Rosie high-performance computing cluster.
- Focuses on a 120-document health-domain corpus.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source provenance or credibility. This work evaluates a simple and interpretable modification to RAG retrieval ranking that incorporates domain-informed source reliability priors. Each document is assigned a prior lambda(s) based on its source type, and retrieval scores are reweighted using score(q, d) = sim(q, d) * lambda(s). The framework is evaluated against a similarity-only baseline on a 120-document health-domain corpus. In this controlled setting, source-aware reranking improves Precision@5 from 0.48 to 0.72 and reduces average adversarial document retrieval under the evaluated threat model, where low-credibility sources are identifiable via metadata. All experiments were executed on Rosie, the high-performance computing cluster at the Milwaukee School of Engineering, which provided the GPU-accelerated infrastructure necessary to run the full experimental pipeline reliably and reproducibly. These results suggest a potential mitigation strategy for source quality degradation in RAG pipelines, within the limits of the experimental setup described.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.22584 [cs.AI] |
| (or arXiv:2607.22584v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22584 arXiv-issued DOI via DataCite |
Submission history
From: Hugo Garrido-Lestache Belinchon [view email]
[v1]
Mon, 8 Jun 2026 14:09:04 UTC (1,032 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.