腾讯开源 AngelSpec 框架:破解大模型真实推理效率难题
Quick Answer
Tencent has open-sourced the AngelSpec framework to enhance inference throughput for large models, addressing the high costs of autoregressive decoding.
Quick Take
The DFly architecture significantly improves performance, achieving the highest average throughput in tests with concurrent requests ranging from 4 to 64, thus providing robust support for efficient model deployment.
Key Points
- AngelSpec framework enhances inference throughput for large models.
- DFly architecture optimizes feature utilization and block dependency modeling.
- Achieved significant throughput improvements over traditional autoregressive decoding.
- Dynamic validation adjusts verification depth based on online load and hardware.
- Open-source project available on GitHub for community use.
📖 Reader Mode
~1 min read在当前大模型规模与服务需求持续增长的背景下,如何降低自回归解码的高昂成本成为了行业的一大痛点。为此,腾讯技术团队正式开源了统训框架
针对不同应用场景下文本特征的差异,

在架构设计上,该框架提出了名为 DFly 的创新块扩散框架。它通过混合目标条件编码骨干与前序条件自回归头,在保持高吞吐量并行生成的同时,大幅优化了目标特征利用率与块内依赖建模。此外,系统引入了动态验证机制,可结合当前在线负载、硬件条件及前缀置信度,在运行时自适应调整验证深度,从而在计算资源与吞吐效率之间实现平衡。

在 Hy3模型系列的系统测试中,搭载 DFly 的方案展现出显著的性能优势。在并发数从4到64的各项测试中,该方案均实现了最高的平均吞吐量,相比传统的自回归解码提升显著,为大模型的工程化落地与高效推理提供了坚实的开源工具支持。
项目地址:https://github.com/Tencent/AngelSpec
论文:https://arxiv.org/pdf/2607.25852
— Originally published at news.aibase.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.