LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
Quick Answer
This paper shows that LISA (Linear-Indexed Sparse Attention) enhances long-context reasoning by reducing inference complexity from O(n^2) to O(nM) using a Linear Attention module and a Lightning Indexer.
Quick Take
Experiments show a 50% speedup in inference with 16K-token contexts and a 5.6% performance improvement on benchmarks like AIME and MATH-500.
Key Points
- LISA integrates Linear Attention and Lightning Indexer for efficient long-context processing.
- Reduces inference complexity significantly, enabling better scalability in production settings.
- Achieves 50% faster inference speeds on 16K-token contexts.
- Improves reasoning performance by 5.6% on benchmarks like AIME and MATH-500.
- Utilizes a two-stage training pipeline for optimal model integration.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths under the test-time scaling paradigm. However, the O(n^2) computational complexity of standard self-attention causes inference costs to grow sharply with long sequences, limiting the deployment of long-CoT reasoning in production settings. To address this, we propose LISA (Linear-Indexed Sparse Attention), a plug-and-play attention replacement module that requires no pretraining from scratch. LISA integrates two lightweight components in parallel within the original model: (1) a Linear Attention module that provides long-range memory with O(n) time complexity; (2) a Lightning Indexer that selects the top-M important tokens from the full context to feed into a Sparse Self-Attention. The two branches are fused via a gating mechanism, reducing inference complexity from O(n^2) to O(nM) (M << n) for generating n tokens. We design a two-stage training pipeline: Stage 1 initializes the model by integrating the linear attention to capture long-range dependencies, complemented by a sliding-window attention mechanism that is optimized via knowledge distillation to approximate the full self-attention distribution of a frozen teacher model. In Stage 2, we further introduce the Indexer to replace the static sliding-window mechanism, enabling dynamic token selection from broader contexts. The Indexer is trained using a novel per-head KL divergence loss, which aligns its selection behavior with the attention patterns of the teacher model. Experiments on DeepSeek-distilled-Qwen models demonstrate that LISA achieves a 50% inference speedup under 16K-token context, while improving average performance by 5.6% on reasoning benchmarks including AIME and MATH-500.
| Comments: | 20 pages, 10 figures |
| Subjects: | Artificial Intelligence (cs.AI) |
| Report number: | LISA-v2-2026 |
| Cite as: | arXiv:2607.19358 [cs.AI] |
| (or arXiv:2607.19358v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19358 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhao Yu [view email]
[v1]
Fri, 29 May 2026 03:05:29 UTC (161 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.