Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking
Quick Answer
This paper shows that The authors introduce a novel evaluation framework, SetwiseEvalKit, for document set selection and ranking, addressing inter-document interactions.
Quick Take
Their method, Rubric4Setwise, outperforms existing rerankers by achieving superior downstream generation performance with fewer documents, maintaining state-of-the-art results across various scenarios.
Key Points
- SetwiseEvalKit includes 28K high-quality evaluation rubrics for diverse document scenarios.
- 12 rerankers were evaluated, with the best achieving only 45% coverage.
- Rubric4Setwise is training-free and converts rubric-based criteria into selection signals.
- The method shows improved performance with fewer documents and search rounds.
- It effectively closes the loop from evaluation to optimization in document retrieval.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Kailin Jiang, Lei Liu, Jian Xi, Hui Xu, Junlin Liu, Baochen Fu, Shaoqing Ren, Bin Li, Vichwang, Yu Lu, Haibo Shi
Abstract:As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.
| Comments: | Project Page: this https URL |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.19747 [cs.CL] |
| (or arXiv:2607.19747v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19747 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kailin Jiang [view email]
[v1]
Wed, 22 Jul 2026 04:45:10 UTC (16,566 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.