LRCC: Generalizing Low-Rank Compression with Conditional Computation
Quick Answer
This paper shows that The Low-Rank Conditional Computation (LRCC) method enhances pretrained language models like Llama and Qwen by introducing token-dependent computation, achieving a 7.6 percentage-point increase in average downstream accuracy on Llama-2-7B compared to static low-rank compression.
Quick Take
LRCC optimizes lightweight routers while keeping low-rank factors frozen, improving both perplexity and accuracy without specialized kernels.
Key Points
- LRCC introduces token-dependent computation for pretrained models, enhancing efficiency.
- Achieves a 7.6 percentage-point accuracy gain on Llama-2-7B over static methods.
- Improves perplexity and accuracy on Llama-3.2-1B at matched decoding latency.
- Evaluated on Llama and Qwen models for language modeling and zero-shot tasks.
- Lightweight routers are optimized while low-rank factors remain frozen during training.
DeepSignal Analysis
What happened
The Low-Rank Conditional Computation (LRCC) method introduces token-dependent computation to pretrained language models, specifically enhancing models like Llama and Qwen. This approach resulted in a 7.6 percentage-point increase in average downstream accuracy on Llama-2-7B compared to static low-rank compression methods.
Key evidence
- LRCC optimizes lightweight routers while keeping low-rank factors frozen, allowing for improved performance without altering the underlying model structure.
- The method was evaluated on Llama and Qwen models, demonstrating enhanced predictive performance in both language modeling and zero-shot downstream tasks.
- At matched batch-size-1 decoding latency, LRCC showed improvements in both perplexity and downstream accuracy on Llama-3.2-1B, maintaining competitiveness on Llama-2-7B.
Why it matters
The introduction of LRCC could signify a shift in how pretrained models are optimized, potentially leading to more efficient computation in language models. By allowing for token-dependent computation, LRCC may enhance model adaptability and performance across various tasks, which is crucial for applications requiring real-time processing and accuracy.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves the predictive performance over static low-rank compression, including a 7.6 percentage-point gain in average downstream accuracy on Llama-2-7B over static methods. At matched batch-size-1 decoding latency, LRCC improves both perplexity and downstream accuracy on Llama-3.2-1B and remains competitive on Llama-2-7B, without specialized kernels. Finally, we assess the usefulness of assigning a token-wise path by analyzing the routers' path choices.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08858 [cs.CL] |
| (or arXiv:2610.08858v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08858 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Thomas Vaitses Fontanari [view email]
[v1]
Mon, 5 Oct 2026 09:01:51 UTC (335 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.