MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning
Quick Answer
MILES introduces a dynamic memory framework for large language models that enhances reasoning by learning selection policies for modular memory units.
Quick Take
It outperforms existing methods in accuracy and efficiency, demonstrating robust performance across various tasks with limited supervision.
Key Points
- MILES uses modular memory units with learnable selection heads for improved reasoning.
- The framework supports coarse-to-fine retrieval for better memory utilization.
- Extensive experiments show MILES achieves superior accuracy-efficiency tradeoffs.
- MILES is designed for incremental memory expansion under test-time constraints.
- It consistently matches or outperforms prior memory-based methods.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not optimized for final-answer correctness. Learning selection policies requires large-scale training data and fixed action spaces, making such approaches unsuitable for test-time settings where memory expands incrementally and only limited supervision is available. We propose MILES (Modular Instruction Memory with LEarnable Selection for self-improving LLM reasoning), a framework that dynamically expands step-wise memory and applies correctness-optimized memory composition under realistic test-time constraints. MILES maintains modular memory units consisting of asymmetric pairs of sub-goal embeddings and sub-instructions, each associated with a learnable selection head. This memory structure enables a coarse-to-fine retrieval mechanism: The coarse level enables memory expansion and collects supervision for training selection heads from confident samples, while the fine stage applies learned selection heads to rerank coarse-level candidates and guide reasoning for uncertain samples. MILES consistently matches or outperforms prior methods while achieving superior accuracy-efficiency tradeoffs. Extensive experiments demonstrate its effectiveness, robustness, and transferability.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.06974 [cs.CL] |
| (or arXiv:2607.06974v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.06974 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruilin Tong [view email]
[v1]
Wed, 8 Jul 2026 03:51:37 UTC (1,089 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.