How Together AI built the world’s fastest speech-to-text stack
Quick Answer
Together AI developed the fastest speech-to-text stack, achieving 20 hours of transcription in under 10 seconds using NVIDIA's Parakeet-TDT 0.6B v3 and OpenAI's Whisper Large v3.
Quick Take
Key optimizations included TensorRT for encoder efficiency and GPU-based decoding, resulting in a 2-3x faster performance.
Key Points
- NVIDIA Parakeet-TDT 0.6B v3 transcribes 20 hours of speech in under 10 seconds.
- TensorRT optimization improved encoder efficiency and reduced memory usage.
- GPU-based decoding eliminated CPU bottlenecks, achieving 2-3x faster performance.
- Audio preprocessing was streamlined to reduce latency and improve throughput.
- The stack supports both offline and streaming transcription modes.
Source Excerpt
Together AI built the fastest speech-to-text stack on Artificial Analysis by treating ASR as a full-path systems problem, not just a GPU inference problem.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.




