Inside the Together AI kernels team
Quick Answer
The Together AI kernels team, led by Dan Fu and Tri Dao, achieved 2-3x speedups in GPU performance with FlashAttention, revolutionizing AI-native cloud infrastructure.
Quick Take
Their ThunderKittens library enabled rapid adaptation to NVIDIA's Blackwell GPUs, producing FP4 and FP8 GEMM kernels with up to 2x speed improvements over cuBLAS.
Key Points
- FlashAttention challenged existing GPU optimization limits, achieving significant performance gains.
- ThunderKittens reduced kernel code complexity from over 1,000 lines to 100-200 lines.
- The team produced some of the fastest FP4 and FP8 GEMM kernels for Blackwell within a week.
- Collaboration with academic institutions fuels innovative kernel development at Together AI.
- The gap between model design and hardware efficiency is crucial for AI-native applications.
Source Excerpt
The team behind FlashAttention and ThunderKittens — how Together AI's kernel researchers close the gap between GPU hardware and production AI.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Together AI
See more →
Open, convenient and predictable: Introducing Provisioned Throughput
Together AI introduces Provisioned Throughput, offering guaranteed inference capacity for MiniMax M3 and GLM-5.2 at $0.05 per PTU per minute, achieving costs up to 90% lower than Claude Opus 4.8. This new model provides predictable pricing and a 99% uptime SLA, catering to companies transitioning to open weight models for production workloads.




