
Autoscaling endpoints for LLM inference
Quick Answer
Together AI introduces autoscaling for LLM inference, allowing deployments to scale based on metrics like in-flight requests and GPU utilization, optimizing performance under peak traffic.
Quick Take
The platform enables users to set replica bounds and select metrics, addressing the challenges of over- and under-provisioning while ensuring low latency for users.
Key Points
- Autoscaling based on metrics like in-flight requests and TTFT enhances deployment efficiency.
- Over-provisioning leads to wasted GPU resources, while under-provisioning causes significant latency spikes.
- The platform provides a catalog of inference-native metrics for tailored autoscaling policies.
- Timing windows for scaling up and down help manage costs and user-facing latency effectively.
- Choosing the right scaling metric is crucial for maintaining optimal performance during traffic peaks.
DeepSignal Analysis
What happened
Together AI has introduced an autoscaling feature for LLM inference that adjusts deployments based on metrics like in-flight requests and GPU utilization. Users can set replica bounds and select metrics to optimize performance during peak traffic. This approach aims to mitigate the costs associated with over- and under-provisioning while maintaining low latency.
Key evidence
- The autoscaling feature allows deployments to scale based on metrics understood by the inference engine, including in-flight requests and GPU utilization.
- Users can set replica bounds and choose metrics, which helps address over- and under-provisioning issues that can lead to increased costs.
- The platform's control loop continuously evaluates observed metrics to determine the desired number of replicas, ensuring that capacity aligns with traffic demands.
Why it matters
This development is significant as it addresses the challenges of managing LLM deployments in a cost-effective manner. By allowing for dynamic scaling based on real-time metrics, organizations can avoid the pitfalls of both over-provisioning, which incurs unnecessary costs, and under-provisioning, which can lead to degraded performance and user experience. The ability to set specific scaling metrics and bounds provides users with greater control over their deployments.
Source Excerpt
GPU utilization can read healthy while your queue backs up, and a new replica takes minutes to warm. Here's how to pick autoscaling metrics, tune scale-up/down windows, and budget for cold starts on dedicated inference.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.




