Post
Quick Answer
AMD has launched the Instella-MoE-16B-A3B, a fully open Mixture-of-Experts model with 16 billion parameters, achieving a benchmark score of 76.7, outperforming other open models.
Quick Take
The model boasts enhanced efficiency through expert-parallel communication and extensive pre-training on AMD's MI300X and MI325X GPUs.
Key Points
- Model features 16B total parameters with 2.8B active per token across 27 layers.
- Achieves 12.7% faster pre-training and 39.2% lower time to first token.
- Scores 76.7, leading among fully open models, ahead of Moonlight-16B-A3B.
- Includes weights from various training stages and research collaborations.
- Available on GitHub and Hugging Face for further exploration.
📖 Reader Mode
~1 min readAMD released Instella-MoE-16B-A3B, a fully open Mixture-of-Experts model trained from scratch on AMD Instinct MI300X and MI325X GPUs. Here are some key takeaways: 1. The model → 16B total parameters, 2.8B active per token → 2 shared experts plus 6 of 64 routed, across 27 layers → 7.1T pre-training tokens, context extended from 4K to 64K 2. Where the efficiency comes from → FarSkip-Collective overlaps expert-parallel communication with computation → 12.7% faster pre-training → Up to 39.2% lower time to first token under expert parallelism in SGLang 3. What it scores → Base averages 76.7, strongest among fully open models → Ahead of Moonlight-16B-A3B at 76.2 and OLMo-3-7B at 70.1 → Behind Qwen3.5-4B-Base at 79.5, which is not fully open 4. What actually ships → Weights from pre-training, mid-training, long-context, SFT, DPO and RL → Data mixtures, training configs and inference code → ResearchRAIL on weights, MIT on the training codebase Full analysis: marktechpost.com/2026/08/01/amd… Model weight: huggingface.co/collections/am… Repo: github.com/AMD-AGI/Instel…
@AMD— Originally published at x.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from WebSearch (Tavily)
See more →全球AI芯片峰会,9月上海见!
The 2026 Global AI Chip Summit will take place in Shanghai on September 22-23, focusing on the evolving AI chip landscape, including the shift from training to inference, the rise of diverse chip technologies, and the restructuring of industry competition. Notable speakers include experts from leading universities and companies, discussing advancements in AI chip architecture and applications.