Tencent Open Sources AngelSpec Framework - AI NEWS
Quick Answer
Tencent has open-sourced the AngelSpec framework to enhance inference throughput for large models, utilizing a novel block diffusion architecture called DFly.
Quick Take
The framework shows significant performance improvements, achieving the highest average throughput in tests with concurrency from 4 to 64, thereby supporting efficient deployment and inference of large models.
Key Points
- AngelSpec targets high-cost autoregressive decoding in large models.
- Utilizes a multi-token prediction mechanism for dialogue tasks.
- Introduces DFly architecture for optimized feature utilization and parallel generation.
- Achieved highest average throughput in tests with concurrency from 4 to 64.
- Supports efficient inference and engineering deployment of large models.
DeepSignal Analysis
What happened
Tencent has released the AngelSpec framework as open-source, designed to improve the inference throughput of large models. It employs a block diffusion architecture called DFly, which optimizes feature utilization and dependency modeling. The framework reportedly achieves the highest average throughput in tests with concurrency levels from 4 to 64.
Key evidence
- The AngelSpec framework aims to enhance inference throughput for large models by addressing the high costs associated with autoregressive decoding.
- AngelSpec utilizes a block diffusion model called DFly, which combines a hybrid target conditional encoding backbone with a previous conditional autoregressive head.
- In tests with the Hy3 model series, the DFly-equipped solution achieved the highest average throughput across concurrency levels ranging from 4 to 64.
Why it matters
The open-sourcing of AngelSpec could significantly impact the deployment of large AI models by providing tools that enhance efficiency and reduce costs. As the demand for large models grows, frameworks like AngelSpec that improve throughput can help organizations manage resources better and optimize performance in real-world applications. This move may also encourage further collaboration and innovation in the AI community.
📖 Reader Mode
~2 min readWith the continuous growth of large model scale and service demand, reducing the high cost of autoregressive decoding has become a major pain point in the industry. To address this, the Tencent technical team officially open-sourced the unified training framework
Considering the differences in text characteristics under different application scenarios,

In terms of architecture design, the framework introduces an innovative block diffusion framework called DFly. It optimizes target feature utilization and intra-block dependency modeling while maintaining high-throughput parallel generation by combining a hybrid target conditional encoding backbone with a previous conditional autoregressive head. Additionally, the system incorporates a dynamic verification mechanism that adaptsively adjusts the verification depth at runtime based on current online load, hardware conditions, and prefix confidence, achieving a balance between computational resources and throughput efficiency.

In system testing of the Hy3 model series, the solution equipped with DFly demonstrated significant performance advantages. In all tests with concurrency ranging from 4 to 64, this solution achieved the highest average throughput, showing significant improvements over traditional autoregressive decoding, providing solid open-source tools support for the engineering deployment and efficient inference of large models.
Project: https://github.com/Tencent/AngelSpec
Paper: https://arxiv.org/pdf/2607.25852
— Originally published at news.aibase.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.