DeepSeek官方开源昇腾基础组件,与昇腾共建高效易用的AI芯片软件生态 – 量子位
Quick Answer
DeepSeek has open-sourced foundational components for Huawei's Ascend platform, including high-performance libraries like DeepGEMM and DeepEP, marking a significant milestone for AI software ecosystems in China.
Quick Take
The release enables deep optimization of memory access and computation pipelines, achieving communication performance close to hardware limits with metrics like 375 GB/s bandwidth.
Key Points
- DeepSeek released components like DeepGEMM and DeepEP for Huawei's Ascend platform.
- The open-source libraries support high-performance operator development and distributed communication.
- Communication performance metrics reached 375 GB/s and 347 GB/s for different operations.
- The infrastructure enables low-latency inference and large-scale model training.
- Huawei's collaboration with DeepSeek aims to enhance AI software ecosystem efficiency.
📖 Reader Mode
~3 min read< img id="wx_img" src="https://www.qbitai.com/wp-content/uploads/imgs/qbitai-logo-1.png" width="400" height="400">
2026-09-30 10:53:17 来源:量子位
9月30日,DeepSeek 正式开源面向昇腾算力平台的基础设施组件,涵盖 TileLang 高级语言编译工具、高性能计算库与分布式通信库,与此前面向GPU平台开源的组件一一对应。Deepseek团队面向华为昇腾平台开源发布基础设施加速库,对中国 AI 软件生态具有里程碑意义,为算力平台提供了新的选择。
本次开源,DeepSeek发布了DeepGEMM、FlashMLA、TileKernel、DeepSelect等高性能昇腾算子库组件和DeepEP分布式通信库。昇腾提供了稳定、开放的Ascend C API接口,使能用户对访存路径及计算流水线进行深度调优,完成高性能算子开发。同时依托开放 Ascend C 编程接口与 PTO ISA 底层指令体系,支持了DeepSeek的TileLang编程开源工作。昇腾生态兼顾不同开发范式,既支持资深开发者进行精细的手工调优,也支持 TileLang 等高级语言的编译对接。
华为向DeepSeek团队提供了联合定义的昇腾超节点SuperPoD Flex 和UBL128组网方案,可实现128卡3.2Tbps单层交换Scale-up网络和256K卡两层交换Scale-out网络,可以支持超低时延推理和前沿基座模型的大规模训练的需求。基于全互联的UBL128超节点,昇腾提供了 ASC-COMM 高性能自定义通信编程库,支撑用户通信算子与通算融合算子高性能编程。Deepseek团队开发了高性能 DeepEP 通信库,涵盖EP / CP / PP / FSDP 等模式下的通信算子,为更大模型向更大集群扩展提供支持,互联带宽实测达到 Dispatch 375 GB/s、Combine 347 GB/s 的通信性能,接近硬件上限。
为支持广泛的开发者基于昇腾 950 及超节点集群部署 DeepSeek模型,华为将此次联合创新的成果在CANN社区开源,包括大EP低延时推理部署、单卡/单机部署、大规模训练、超长文本KVcache池化与 Agentic RL。其中面向大规模推理场景,基于EP32部署策略,offline推理模式下DeepSeek-V4.1-Flash不带框架纯模型性能可实现TPOT=5ms,每卡输出吞吐2469 tokens/s;TPOT=10ms,每卡输出吞吐5102 tokens/s。(注:Benchmark数据均基于Offline推理模式采集,不包含Serving调度和框架负载均衡影响。ContextLength = 128K,Dspark 投机接受率 0.85,profiling链接详见附录:DeepSeek-V4.1-Flash超节点部署推理优化实践)
基承其本,算展其用。华为昇腾平台将与Deepseek等前沿模型团队,持续深耕芯模协同,联合创新。以AI算力底座承载模型能力,深度适配,双向奔赴,打通推理到训练的全链路,携手打造开放、高效、易用的AI软件生态。
附开源参考实践和技术报告链接:
- 通信编程体系:基于UB memory和URMA通信编程接口的通算并行算子开发实践:
https://gitcode.com/cann/cann-recipes-infer/tree/master/docs/models/deepseek_v4_1/deepseek_v4.1_asc_comm_tech_report.md
- 基于Agent-Friendly算子编程范式CANNBot-DSL的高性能融合算子开发实践:
https://gitcode.com/cann/cann-recipes-infer/blob/master/docs/models/deepseek_v4_1/deepseek_v4.1_cannbotdsl_operator_guide.md
- DeepSeek-V4.1-Flash超节点部署推理优化实践:
https://gitcode.com/cann/cann-recipes-infer/tree/master/docs/models/deepseek_v4_1/deepseek_v4.1_flash_cann_tech_report.md
- DeepSeek-V4.1-Flash单机部署低时延推理优化实践:
https://gitcode.com/cann/cann-recipes-infer/tree/master/docs/models/deepseek_v4_1/deepseek_v4.1_low_latency_tp_guide.md
- DeepSeek-V4.1-Flash 单卡部署推理优化实践:https://gitcode.com/cann/cann-recipes-infer/blob/master/docs/models/deepseek_v4_1/dsv41-flash-950-single-card.md
- 基于PyTorch原生框架的DeepSeek-V4.1-Flash模型预训练实践:
https://gitcode.com/cann/cann-recipes-train/blob/master/docs/llm_pretrain/ascend950_superpod_moe_pretrain_system_reference.md
- DeepSeek-V4.1-Flash模型低精度量化训练实践:https://gitcode.com/cann/cann-recipes-train/tree/master/docs/llm_pretrain/ascend950/low_precision_training.md
- DeepSeek-V4 DSpark草稿模型昇腾训练实践:https://gitcode.com/cann/cann-recipes-train/tree/master/llm_sft/DeepSpec
- DeepSeek-V4 单机LoRA 微调实践:https://gitcode.com/cann/torchtitan-npu/blob/master/docs/feature_guides/deepseek_v4_lora.md
- 基于鲲鹏CPU、昇腾950、灵衢协同的Agentic RL训练实践:https://gitcode.com/cann/cann-recipes-train/tree/master/docs/llm_rl/kunpeng_ascend_agentic_rl_infrastructure.md
- 鲲鹏CPU对AgentENV的适配与优化实践:
https://gitcode.com/KunpengSDK/agentENV/blob/main/agentenv-kunpeng-optimization-practice.md
- 推理模型优化Agent 部署调优实践
https://gitcode.com/cann/cann-recipes-infer/blob/master/docs/agent/DeepSeekV4.1-Flash-Model-Agent.md
- AMCT 低精度量化Agent量化实践:
https://gitcode.com/cann/amct/tree/master/examples/models/deepseekv4.1/DeepSeekV4.1-Flash-Quantization-Agent.md
- TileLang 算子Agent开发和优化实践:
https://gitcode.com/cann/cannbot-skills/tree/master/plugins-official/tilelang-op-orchestrator
本文由华为提供,量子位获授权转载,观点归原作者所有。
版权所有,未经授权不得以任何形式转载及使用,违者必究。
— Originally published at qbitai.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.