SpecLA: Efficient Speculative Decoding for Linear-Attention Models
Quick Answer
SpecLA introduces an efficient speculative decoding runtime for stateful linear-attention models, achieving up to 1.70x speedup over traditional autoregressive decoding on NVIDIA H100 with GDN-1.3B target.
Quick Take
It employs topology-aware kernels and confidence pruning to optimize the verification of token candidates, enhancing performance in language tasks.
Key Points
- SpecLA optimizes speculative decoding for linear-attention models, enhancing efficiency.
- Achieves up to 1.70x speedup on NVIDIA H100 with GDN-1.3B benchmark.
- Utilizes topology-aware kernels for verifying token chains and trees.
- Implements confidence pruning to filter useful candidates for verification.
- Addresses limitations of existing speculative systems designed for Transformer KV caches.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Linear-attention models replace the growing KV cache with recurrent states, but autoregressive decoding still reads, updates, and writes these states one token at a time. Speculative decoding can reduce this cost by verifying several draft tokens in one target pass, yet existing speculative systems are designed for Transformer KV caches. For stateful linear-attention targets, verification must follow recurrent dependencies across chains and branches, acceptance must update only the accepted state trajectory, and the drafter must avoid submitting candidates that waste stateful verification work. This paper presents SpecLA, a speculative decoding runtime for stateful linear-attention models. SpecLA verifies chains and trees with topology-aware kernels, stores compact factors produced during verification to recover accepted states, and uses confidence pruning plus a target-aligned EAGLE-style drafter to feed useful candidates to the verifier. On an NVIDIA H100 with a public GDN-1.3B target, SpecLA achieves up to 1.70x end-to-end speedup over autoregressive decoding.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.16673 [cs.CL] |
| (or arXiv:2607.16673v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16673 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Fuliang Liu [view email]
[v1]
Sat, 18 Jul 2026 07:08:19 UTC (6,576 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.