CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
Quick Answer
CMT-RAG introduces a novel memory framework for multi-turn multi-hop retrieval-augmented generation (RAG), enhancing conversational context tracking through structured reasoning traces.
Quick Take
It outperforms five baselines in answer accuracy on the MuMu-QA benchmark, demonstrating improved efficiency in recovering prior reasoning and evidence.
Key Points
- CMT-RAG employs a state-space trace generator for real-time memory incorporation.
- It decomposes queries into structured drafts with retrieval-oriented sub-questions.
- Persistent memory traces are stored in a session-level directed acyclic graph (DAG).
- CMT-RAG shows consistent performance improvements over existing RAG systems.
- MuMu-QA benchmark includes explicit cross-turn sub-question dependency annotations.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG systems typically represent conversational memory as raw dialogue history, rewritten queries, or unstructured summaries, making it difficult to recover the specific prior reasoning steps and evidence required for follow-up queries. Our key insight is to align conversational memory with retrieval by representing dialogue context as sub-question-level reasoning traces. Building on this insight, we introduce MuMu-QA, a benchmark for multi-turn multi-hop RAG with explicit cross-turn sub-question dependency annotations, and CMT-RAG, a complementary memory framework for this setting. At each turn, CMT-RAG employs a state-space trace generator, whose recurrent state serves as runtime memory, to incorporate recent conversational context and decompose the current query into structured trace drafts containing retrieval-oriented sub-questions and dependencies on earlier traces. It then grounds these drafts with retrieved evidence and stores them as persistent memory traces in a session-level DAG, enabling future turns to efficiently recover relevant prior reasoning and evidence. Experiments on MuMu-QA and corpus-level RAG benchmarks show that CMT-RAG consistently outperforms five categories of RAG baselines in answer accuracy.
| Subjects: | Computation and Language (cs.CL); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2607.26470 [cs.CL] |
| (or arXiv:2607.26470v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26470 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Lang Zhou [view email]
[v1]
Wed, 29 Jul 2026 04:50:41 UTC (927 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.