NVIDIA AI on X: "SGLang is hitting 180 tok/s/GPU on DeepSeek-V4 decode with ~1M context on Blackwell. Good to see fast progress in open source DeepSeek-V4 inference on new hardware. This comes from Blackwell-specific optimizations by @lmsysorg that better use the model’s hybrid sparse" / X
Quick Answer
NVIDIA's SGLang achieves 180 tok/s/GPU on DeepSeek-V4 decoding with ~1M context on Blackwell, showcasing significant advancements in open-source inference.
Quick Take
Optimizations by @lmsysorg enhance the model's hybrid sparse attention capabilities, ensuring robust performance from launch.
Key Points
- SGLang reaches 180 tok/s/GPU on DeepSeek-V4 with ~1M context.
- Optimizations by @lmsysorg improve hybrid sparse attention on Blackwell.
- DeepSeek V4 launched with full stack optimizations from architecture to kernels.
- Verified RL training pipeline available for V4 at launch.
📖 Reader Mode
~1 min readPost
Post
SGLang is hitting 180 tok/s/GPU on DeepSeek-V4 decode with ~1M context on Blackwell. Good to see fast progress in open source DeepSeek-V4 inference on new hardware. This comes from Blackwell-specific optimizations by
@lmsysorgthat better use the model’s hybrid sparse attention.
DeepSeek V4 by
@deepseek_aijust dropped! SGLang is ready on Day 0 with a full stack of optimizations from architectures to low-level kernels. We also deliver a verified RL training pipeline in Miles (by
@radixark) for V4 at launch: 1️⃣ Native "ShadowRadix" Design: DeepSeek V4's
Don't miss what's happening
People on X are the first to know.
— Originally published at x.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.

