
Deploying Kimi K3 on Amazon SageMaker HyperPod and Amazon EKS
Quick Answer
Moonshot AI's Kimi K3, a 2.8 trillion parameter Mixture of Experts model, can be deployed on AWS using Amazon SageMaker HyperPod or Amazon EKS.
Quick Take
It offers advanced capabilities for complex tasks and requires substantial GPU resources, specifically the p6-b300 instance for optimal performance.
Key Points
- Kimi K3 features 2.8 trillion parameters with 896 experts, activating 16 per token.
- The model achieves a 2.5x scaling efficiency improvement over its predecessor, Kimi K2.
- Open weights are available on Hugging Face in MXFP4 format for efficient inference.
- Deployment requires p6-b300 instances for optimal GPU compute and performance.
- AWS offers Flexible Training Plans and Capacity Blocks for resource procurement.
DeepSignal Analysis
What happened
Moonshot AI launched Kimi K3, a 2.8 trillion parameter Mixture of Experts model, on July 27, 2026. It can be deployed on AWS using Amazon SageMaker HyperPod or Amazon EKS, requiring substantial GPU resources, specifically the p6-b300 instance for optimal performance.
Key evidence
- Kimi K3 is the first open-weight model to reach the 2.8 trillion parameter class, with 896 specialist experts activating only 16 per token.
- The model supports advanced capabilities such as long-horizon coding and complex reasoning, making it suitable for multi-step workflows.
- Deployment requires a vLLM day-0 inference container and a p6-b300 instance, which includes 8 NVIDIA B300 GPUs for efficient tensor-parallel inference.
Why it matters
The introduction of Kimi K3 signifies a notable advancement in AI model capabilities, particularly in handling complex tasks. Its open-weight availability allows organizations to leverage cutting-edge technology without relying solely on proprietary systems. The infrastructure requirements highlight the increasing demand for high-performance computing resources in AI deployments.
What to watch
Source Excerpt
This post walks through deploying Kimi K3 on AWS using two approaches: Amazon SageMaker HyperPod, and Amazon Elastic Kubernetes Service (Amazon EKS) cluster.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from AWS Machine Learning
See more →
Build an explainable next-best-product recommendation system for banking on AWS
AWS presents a deep learning-based Next-Best-Product recommendation system for banks, utilizing Amazon SageMaker and PyTorch to enhance customer product predictions. This architecture leverages a multi-tower neural network for improved accuracy and explainability, addressing the complexities of customer data in financial services.




