
Configuring Dedicated Model Inference
Quick Answer
Together AI's Dedicated Model Inference architecture integrates endpoints, deployments, and configs, enabling capacity-aware traffic routing for efficient model inference.
Quick Take
This setup supports advanced features like A/B testing and zero-downtime changes, ensuring optimal resource utilization and traffic management.
Key Points
- Endpoints serve as stable identifiers for model inference requests.
- Deployments link specific models to configs with autoscaling policies.
- Traffic routing is based on deployment weights and ready replicas.
- Config profiles optimize for GPU type, count, and performance metrics.
- New deployments require traffic split configuration to receive requests.
DeepSignal Analysis
What happened
Together AI has introduced a Dedicated Model Inference architecture that consists of endpoints, deployments, and configs. This architecture allows for capacity-aware traffic routing, enabling features like A/B testing and zero-downtime changes. The system is designed to optimize resource utilization and manage traffic effectively.
Key evidence
- The architecture includes three components: endpoints, deployments, and configs, which work together for efficient model inference.
- Traffic routing is based on a weight system that considers the number of ready replicas, allowing for proportional traffic distribution.
- Each deployment is linked to a specific model and config, with the ability to scale and absorb traffic based on its capacity.
Why it matters
This architecture is significant as it enhances the efficiency of model inference by allowing for dynamic traffic management and resource allocation. The ability to perform A/B testing and shadow experiments without downtime can lead to better model performance and user experience. Furthermore, the immutable nature of configs ensures stability and reliability in deployments.
What to watch
Source Excerpt
The three-part resource model behind Together AI Dedicated Model Inference—endpoints, deployments, configs—and how capacity-aware routing ties them together.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.

