LoRi: Low-Rank Distillation for Implicit Reasoning
Quick Answer
LoRi introduces a low-rank distillation framework for implicit reasoning in large language models like LLaMA and Qwen, enhancing performance on multi-step tasks.
Quick Take
The method aligns reasoning trajectories in a low-rank tensor subspace, achieving results close to explicit chain-of-thought prompting and outperforming previous iCoT methods across various benchmarks.
Key Points
- LoRi aligns teacher and student reasoning trajectories in a shared low-rank tensor subspace.
- The framework captures global reasoning structure while enabling a compact latent process.
- Evaluated on LLaMA and Qwen, it shows consistent improvement on mathematical reasoning tasks.
- Performance approaches explicit chain-of-thought accuracy, especially on challenging multi-step tasks.
- Outperforms previous implicit chain-of-thought distillation methods across multiple model families.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Implicit chain-of-thought (iCoT) methods aim to internalize reasoning in large language models, but often underperform explicit CoT prompting. We empirically find that hidden-state reasoning trajectories exhibit low-rank structure. Motivated by this observation, we propose a low-rank distillation framework that transfers reasoning by aligning teacher and student trajectories in a shared low-rank tensor subspace using first- and second-order statistics. The resulting formulation captures the global structure of reasoning while supporting a compact latent reasoning process. We evaluate the method across multiple model families, including LLaMA and Qwen, at different scales on mathematical reasoning benchmarks. Our approach consistently improves performance, especially on challenging multi-step tasks, approaching explicit CoT accuracy and outperforming prior iCoT distillation methods.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.05315 [cs.CL] |
| (or arXiv:2606.05315v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.05315 arXiv-issued DOI via DataCite |
Submission history
From: Ryan Solgi [view email]
[v1]
Wed, 3 Jun 2026 18:05:50 UTC (1,259 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.