RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
Quick Answer
RIMS introduces a three-stage preference optimization framework for small-scale language models (SLMs) in retrieval-augmented generation, enhancing performance under noisy conditions.
Quick Take
It outperforms existing methods like RoseRAG on multi-hop question answering benchmarks, achieving significant improvements in Exact Match and F1 scores. The approach leverages synthetic data generation and a differentiable soft aggregation mechanism to optimize preference selection.
Key Points
- RIMS utilizes synthetic chain-of-thought preference data generation for improved SLM performance.
- The framework includes a differentiable soft aggregation mechanism for better gradient alignment.
- Experiments show RIMS outperforms state-of-the-art methods on four multi-hop question answering benchmarks.
- Significant gains in Exact Match and F1 scores were achieved under noisy retrieval conditions.
- Implementation details are available at the provided URL.
DeepSignal Analysis
What happened
RIMS presents a three-stage framework aimed at optimizing preferences for small-scale language models in retrieval-augmented generation scenarios. This method reportedly enhances performance in noisy conditions, surpassing existing models like RoseRAG on multi-hop question answering benchmarks.
Key evidence
- RIMS employs synthetic data generation through rejection sampling using the target small-scale language model, avoiding reliance on proprietary models.
- The framework introduces a differentiable soft aggregation mechanism, which maintains gradient signals from all preference pairs while preserving the structure of margin-aware selection.
- Experiments demonstrate that RIMS achieves significant improvements in Exact Match and F1 scores across four multi-hop question answering benchmarks under noisy retrieval conditions.
Why it matters
The development of RIMS is significant as it addresses the limitations of existing preference-based methods that either discard useful signals or treat comparisons independently. By optimizing preference selection in a more nuanced way, RIMS could enhance the utility of small-scale language models in practical applications, especially in environments where resources are limited and noise is prevalent.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Small-scale language models (SLMs) are attractive for retrieval-augmented generation (RAG) in resource-constrained settings, but their limited capacity makes them highly sensitive to noisy or spurious retrieved evidence. Existing preference-based methods such as RoseRAG select only the hardest single preference pair via hard argmin/argmax, discarding the remaining signal; others treat multiple pairs as independent binary comparisons, resulting in low data utilization. We propose RIMS, a three-stage preference optimization framework comprising (1) synthetic chain-of-thought preference data generation via rejection sampling using the target SLM itself without relying on proprietary models, (2) a differentiable soft aggregation mechanism that replaces hard selection with a smooth operator, preserving gradient signal from all preference pairs while retaining the discriminative structure of margin-aware selection, and (3) preference optimization with the smoothed objective applied to multiple alignment algorithms. We theoretically show that the smoothed approximation admits a controllable error bound and that smooth aggregation yields provably tighter gradient alignment to the oracle objective than hard selection. Experiments on four multi-hop question answering benchmarks show that our approach outperforms state-of-the-art baselines across multiple SLM backbones, achieving consistent gains in Exact Match and F1 under noisy retrieval conditions. Our implementation is available at this https URL.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.16431 [cs.CL] |
| (or arXiv:2607.16431v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16431 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | COLM 2026 |
Submission history
From: Haoyu Wang [view email]
[v1]
Fri, 17 Jul 2026 18:24:19 UTC (1,224 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.