
Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Quick Answer
NVIDIA's GB300 NVL72 achieved a world record of 1,648 TFLOPs per GPU while pre-training the DeepSeek-V3 671B model, showcasing significant advancements in MoE architecture and communication efficiency across GPUs.
Quick Take
This performance boost enables researchers to train larger models and conduct more experiments rapidly on NVIDIA's infrastructure.
Key Points
- DeepSeek-V3 model activates ~37B parameters per token, enhancing computational efficiency.
- NVIDIA NVLink provides 1.8 TB/s bandwidth, allowing GPUs to communicate in a single hop.
- Megatron Core achieved 1,648 TFLOPs per GPU, nearly 3x higher than previous models.
- Software optimizations improved performance by 1.5x in six months on GB300 NVL72.
- NVIDIA actively contributes to open-source frameworks like TorchTitan and JAX for better performance.
DeepSignal Analysis
What happened
NVIDIA's GB300 NVL72 achieved a world record of 1,648 TFLOPs per GPU while pre-training the DeepSeek-V3 671B model. This performance reflects advancements in mixture of experts (MoE) architecture and communication efficiency across GPUs, enabling faster training of larger models.
Key evidence
- The GB300 NVL72 set a record of 1,648 TFLOPs per GPU during the pre-training of the DeepSeek-V3 671B model.
- DeepSeek-V3 has 671 billion parameters but activates only about 37 billion parameters per token, enhancing computational efficiency.
- NVIDIA's software optimizations have improved performance by 1.5 times over six months on the GB300 NVL72 platform.
Why it matters
This achievement illustrates the ongoing evolution of AI training infrastructure, particularly in the context of MoE architectures. As communication efficiency becomes a critical factor in scaling AI models, the ability to train larger models more rapidly can significantly accelerate research and development in AI. The record-setting performance also highlights the importance of integrated hardware and software solutions in achieving these advancements.
What to watch
Source Excerpt
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token…
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from NVIDIA Developer Blog
See more →
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
NVIDIA's NeMo pipeline generates 502,536 unique financial news headlines in 82 iterations, addressing data imbalance in financial NLP. The iterative approach uses semantic deduplication and category-weighted sampling to enhance diversity and relevance in generated content.

