Benchmarking inference at scale: coding agents
Quick Answer
Together AI's new benchmark for coding agents reveals that its Together Inference Engine achieves 31% higher TPS than TensorRT-LLM, maintaining under 1s TTFT at 625 TPM per GPU.
Quick Take
This performance is crucial for handling high concurrency and long context requests in production environments.
Key Points
- Together Inference Engine outperforms TensorRT- with 31% higher TPS.
- TTFT remains under 1 second, crucial for user experience.
- Benchmark simulates high concurrency with long input requests.
- Performance gains achieved through full-stack profiling and optimization.
- EAGLE speculative decoding used for improved efficiency.
Source Excerpt
Real-world inference benchmarks for coding agents: 31% more TPS than TensorRT-, 2× better TTFT at saturation, and 76% lower cost than Claude Opus 4. 6.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.




