SeKV: Resolution-Adaptive KV Cache with Hierarchical Semantic Memory for Long-Context LLM Inference
Quick Answer
SeKV introduces a resolution-adaptive KV cache for long-context LLMs, enhancing semantic memory without information loss.
Quick Take
It achieves a 5.9% performance improvement over existing methods while reducing GPU memory usage by 53.3% at 128K context, with minimal additional parameters.
Key Points
- SeKV organizes context into entropy-guided semantic spans for efficient memory use.
- It retains a lightweight summary vector on GPU for coarse routing during inference.
- The method reduces GPU memory requirements by 53.3% compared to full KV caching.
- SeKV improves semantic compression performance by an average of 5.9% across benchmarks.
- Less than 0.05% additional trainable parameters are needed to implement SeKV.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models increasingly operate over long contexts, where the KV cache becomes a dominant memory bottleneck: its size grows linearly with sequence length and must be retained throughout decoding, making full GPU caching prohibitively expensive without compression. Existing KV cache compression methods struggle to balance efficiency with faithful context preservation. Token eviction discards information, while semantic grouping fixes compression decisions at prefill time; neither can recover token-level detail from a compressed span once it becomes relevant during generation. As a solution, we propose SeKV, a resolution-adaptive semantic KV cache that organizes context into entropy-guided semantic spans and stores them across a GPU-CPU memory hierarchy without discarding information. Each span keeps a lightweight summary vector on GPU for coarse routing and a low-rank SVD basis on CPU for on-demand token-level reconstruction. A trained zoom-in mechanism selectively expands query-relevant spans during decoding, enabling precise retrieval without materializing the full KV cache on GPU. SeKV enables adaptive token-level reconstruction while keeping the base LLM fully frozen and adding fewer than 0.05% trainable parameters. Across four benchmarks, SeKV improves over the strongest semantic compression baseline by 5.9% on average while reducing GPU memory by 53.3% versus full KV caching at 128K context. Code is available on this https URL.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.31145 [cs.CL] |
| (or arXiv:2606.31145v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31145 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Amirhossein Abaskohi [view email]
[v1]
Tue, 30 Jun 2026 05:18:02 UTC (623 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.