Expert Coupling in MoE Pretraining: Reducing All-to-All Overhead with Correlated Placement and Token Shuffling
Quick Answer
The study introduces expert coupling in Mixture-of-Experts (MoE) pretraining, significantly reducing all-to-all communication overhead by leveraging correlated expert placements and token shuffling.
Quick Take
This approach enhances token-expert assignments on the same GPU from 12.5% to 59%, achieving up to 2.63X reduction in all-to-all time and 1.41X faster end-to-end training in Megatron-LM across various expert parallelism degrees.
Key Points
- Correlated expert placement reduces inter-GPU communication by grouping frequently selected experts.
- Token shuffling increases token-expert assignments on the same GPU from 12.5% to 59%.
- Achieves 1.16-2.63X reduction in all-to-all communication time.
- End-to-end training time improved by up to 1.41X in Megatron-LM.
- Methods do not alter routing decisions or expert parameters.
DeepSignal Analysis
What happened
The study presents a method called expert coupling in Mixture-of-Experts (MoE) pretraining, which reduces communication overhead during training. By utilizing correlated expert placements and token shuffling, the method improves token-expert assignments on the same GPU significantly.
Key evidence
- In a setup with 8 AMD Instinct MI300X GPUs per node, all-to-all communication can consume 45% of the training step at expert parallelism degree 32 with top-2 routing.
- The proposed methods increase the share of token-expert assignments served on the token's GPU from 12.5% to 59%, effectively reducing communication across GPUs.
- In Megatron-LM, the techniques achieve a reduction in all-to-all time by up to 2.63X and an end-to-end step time improvement of 1.41X across various expert parallelism degrees.
Why it matters
Reducing communication overhead is crucial for efficient training of large models, particularly in distributed settings. The findings suggest that leveraging correlations in expert assignments can lead to significant performance improvements, which may influence future research and development in MoE architectures. This could enhance the scalability and efficiency of training large-scale AI models, making them more accessible for practical applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Mixture-of-Experts (MoE) layers replace the feed-forward block of a Transformer with E expert networks, and each token is routed to k of these experts. Under expert parallelism (EP) the experts are distributed across GPUs, and every MoE layer runs all-to-all collectives in the forward and backward passes to dispatch tokens to their experts and then combine the results. On a cluster with 8 AMD Instinct MI300X GPUs per node, these collectives can take 45% of the training step at EP32 with top-2 routing and 60% with top-6 routing. We find that early in pretraining routers have already learned to assign tokens to experts in correlated patterns, both within a layer and across layers. At top-2, 0.8% of the expert pairs in a layer are selected together by 42% of tokens, and the experts a token selects at one layer predict the experts it selects at the next layer. We use these correlations to keep more token--expert assignments on the token's own GPU, which reduces communication across GPUs and across nodes. Correlated expert placement puts experts that are often selected together on the same GPU. Combined with a dispatcher that sends each token to each GPU once, it removes up to 58% of dispatched rows. Token shuffling applies when sequence parallelism shards tokens across the EP group. It moves each token to the GPU predicted to hold its next-layer experts during the reduce-scatter that follows attention. On one node this raises the share of token--expert assignments served on the token's GPU from 12.5% to 59%. In Megatron-LM, across EP degrees from 8 to 64 with top-2 and top-6 routing, the two methods reduce all-to-all time by 1.16-2.63X and end-to-end step time by up to 1.41X. Neither method changes the models' underlying routing decisions or expert parameters.
| Subjects: | Computation and Language (cs.CL); Distributed, Parallel, and Cluster Computing (cs.DC) |
| Cite as: | arXiv:2610.09372 [cs.CL] |
| (or arXiv:2610.09372v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09372 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Quentin Anthony [view email]
[v1]
Wed, 7 Oct 2026 03:26:24 UTC (4,516 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.