VarRate: Training-Free Variable-Rate KV Cache Compression for Long-Context LLMs
Quick Answer
VarRate introduces a training-free KV codec for long-context LLMs, achieving only 3.5-5.5 accuracy points degradation compared to traditional methods.
Quick Take
It maintains a variable low-rank budget for tokens, outperforming uniform low-rank coding and matching the performance of existing adaptive codecs without training. On LongBench, VarRate stays within 0.8 points of uncompressed models Llama-3.1-8B and Qwen2.5-7B.
Key Points
- VarRate maintains nonzero rank for all tokens, reducing accuracy loss during query-agnostic reuse.
- It outperforms uniform low-rank coding and matches adaptive codecs without requiring training.
- Achieves a maximum accuracy drop of only 3.5-5.5 points under challenging conditions.
- On LongBench, VarRate remains within 0.8 points of uncompressed models across 16 tasks.
- Significantly better than its uniform-rank ablation and competitive with KVzip.
DeepSignal Analysis
What happened
VarRate is a new training-free key-value codec designed for long-context large language models (LLMs). It allocates a variable low-rank budget to tokens based on their query salience, resulting in minimal accuracy degradation compared to traditional methods. VarRate outperforms uniform low-rank coding and matches the performance of adaptive codecs without requiring training.
Key evidence
- VarRate achieves only 3.5-5.5 accuracy points degradation compared to traditional methods, which is significantly lower than the 11-15 points collapse seen in token-selection methods.
- On LongBench, VarRate maintains performance within 0.8 points of uncompressed models Llama-3.1-8B and Qwen2.5-7B while using a matched 20% budget.
- VarRate demonstrates superior performance over uniform low-rank coding and is accuracy-equivalent to KVzip in three out of four settings, with only one-eighth of the prefill overhead.
Why it matters
The development of VarRate addresses a critical bottleneck in LLM inference related to memory management. By maintaining a variable low-rank budget for tokens, it enhances efficiency without sacrificing accuracy. This could lead to more effective deployment of LLMs in real-world applications, especially where memory constraints are a concern. The ability to achieve competitive performance without training also suggests potential cost savings and faster implementation times for developers.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The key-value (KV) cache is the main memory bottleneck in long-context large language model (LLM) inference. Two leading training-free families are both structurally limited: token-selection methods (SnapKV, Ada-KV) score importance from an observation window and evict low-scoring tokens, but eviction is irreversible -- so when the importance signal degrades under query-agnostic reuse, accuracy collapses by 11-15 points; uniform low-rank coding keeps every token but spends equal rank everywhere, wasting budget. We observe that both failures share one cure: rank should be allocated, not evicted. We present VarRate, a training-free KV codec that assigns each token a variable low-rank budget by its query salience, keeping every token at a nonzero rank. Comparable adaptive-rank codecs reach this allocation only through training; VarRate requires none. Because no token is dropped, it degrades by only 3.5-5.5 points where query-aware selection collapses. At a matched 20% budget on LongBench (16 tasks), VarRate stays within 0.8 points of the uncompressed model on both Llama-3.1-8B and Qwen2.5-7B. Averaged over the two, it is the strongest matched-memory compressor. It significantly beats its uniform-rank ablation on both models. Against KVzip, a method purpose-built for query-agnostic reuse, it is accuracy-equivalent in three of four settings and within a point overall, at about one-eighth the prefill overhead.
| Comments: | 20 pages, 8 figures, 24 tables. Includes appendix with additional experiments and analyses |
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15498 [cs.CL] |
| (or arXiv:2607.15498v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15498 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shahrzad Esmat [view email]
[v1]
Thu, 16 Jul 2026 23:03:14 UTC (416 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.