MemDefrag: Latent Memory Defragmentation for Large Language Models
Quick Answer
MemDefrag introduces a training-free, model-agnostic framework for latent memory defragmentation in large language models, significantly improving knowledge retention (43.0% vs.
Quick Take
17.4%/17.6% after 50 updates) and long-context benchmarks. By utilizing a middle-layer tracing signal, it effectively ranks, reorders, and filters memories, addressing performance degradation during updates.
Key Points
- MemDefrag outperforms MemoryLLM and M+ in knowledge retention and long-context tasks.
- It utilizes a middle-layer tracing signal to enhance memory management.
- The framework is training-free and applicable across various .
- Performance degradation during memory updates is significantly reduced.
- Achieves a 43.0% knowledge retention rate after 50 memory updates.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Latent memory, which stores past knowledge fragments as per-layer hidden states, has emerged as a promising paradigm (e.g., MemoryLLM and M+) for long-term memory in large language models (LLMs). However, the paradigm suffers from significant performance degradation during memory updates, due to positional encoding misalignment and the absence of any tracing mechanism to distinguish target memory fragments from irrelevant ones. To discover such a tracing mechanism, we probe the layer-wise attention density over stored memory fragments, and find that a small set of middle transformer layers consistently concentrates the highest density on the target fragment - exposing an inherent tracing signal. In light of this, we propose MemDefrag, a training-free and model-agnostic framework that (1) uses a middle-layer tracing signal to conduct memory defragmentation (rank, reorder, and filter memories), and (2) applies an informativeness-guided proportional forgetting mechanism once capacity is exceeded. Experiments show that MemDefrag substantially outperforms MemoryLLM and M+ on knowledge retention (e.g., 43.0% vs. 17.4%/17.6% after 50 memory updates) and long-context benchmarks, and generalizes well across various LLMs and latent-memory variants.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.05969 [cs.CL] |
| (or arXiv:2607.05969v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05969 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ruiyi Yan [view email]
[v1]
Tue, 7 Jul 2026 08:04:34 UTC (763 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.