OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
Quick Answer
OnlineQAT introduces a two-stage framework for quantization-aware training of large language models, achieving significant performance improvements.
Quick Take
On Qwen3-1.7B, it outperforms existing methods, achieving 57.28 at W3A16 and 32.52 at W2A16, surpassing ReasoningQAT by 2.90 and 0.44 points respectively. This method leverages student-generated responses for better recovery signals.
Key Points
- OnlineQAT employs block-wise QAT for low-bit initialization followed by on-policy distillation.
- The framework uses a frozen full-precision teacher to guide training with reverse-KL signals.
- Achieved best average performance among quantized methods on Qwen3-1.7B.
- Improvements of 2.90 and 0.44 points over ReasoningQAT at W3A16 and W2A16 respectively.
- Demonstrates that student-visited states enhance recovery beyond fixed-completion training.
DeepSignal Analysis
What happened
OnlineQAT presents a two-stage framework for quantization-aware training of large language models, specifically targeting ultra-low-bit quantization. The method demonstrates improved performance on the Qwen3-1.7B model, surpassing previous techniques like ReasoningQAT.
Key evidence
- OnlineQAT achieves 57.28 at W3A16 and 32.52 at W2A16 on the Qwen3-1.7B model, outperforming ReasoningQAT by 2.90 and 0.44 points respectively.
- The framework consists of a low-bit initialization through block-wise quantization-aware training followed by on-policy distillation using student-generated responses.
- A frozen full-precision teacher model provides a sampled reverse-KL training signal at each visited prefix, enhancing recovery signals.
Why it matters
The advancements in OnlineQAT could significantly enhance the deployment of large language models in resource-constrained environments by maintaining accuracy during aggressive quantization. This is particularly relevant as the industry increasingly seeks efficient models for real-time applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, OnlineQAT obtains the best average among the compared quantized methods: 57.28 at W3A16 and 32.52 at W2A16, im- proving over ReasoningQAT by 2.90 and 0.44 points, respectively. The results suggest that student-visited states provide a useful recovery signal beyond fixed-completion training, particularly at three bits.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09346 [cs.CL] |
| (or arXiv:2610.09346v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09346 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenjun Wang [view email]
[v1]
Wed, 7 Oct 2026 03:05:42 UTC (66 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.