Serving MiniMax-M3 for efficient inference: Unlocking 1M-Token Context and Multimodality Without Regrets
Quick Answer
MiniMax's M3 model introduces a 1M-token context and multimodal capabilities, optimized for efficient inference with a 9x speedup in prefill and 15x in decoding, supported by Together AI's cloud infrastructure.
Key Points
- M3 features MiniMax Sparse Attention, reducing long-context processing costs significantly.
- The model supports multimodal reasoning with integrated image and video processing capabilities.
- Together AI collaborated with MiniMax to address engineering challenges for 1M context length.
- Optimizations include KV-Block-Major Sparse Attention and integration with Paged Attention.
- The new architecture enhances throughput by 5% during decoding.
Source Excerpt
How Together served MiniMax-M3 efficiently with KV-block-major sparse attention, paged MSA decode, optimized index scoring, and a Rust-based multimodal gateway.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

