
The Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Quick Answer
Baseten has raised $13B, emerging as a leading AI infrastructure decacorn, while Philip Kiely and Ali Taha discuss the evolution of inference engineering, highlighting how quantization can enhance throughput by 20% without sacrificing model quality.
Quick Take
Their insights reveal that inference is now a distinct engineering discipline critical for optimizing AI models.
Key Points
- Baseten's valuation skyrocketed to $13B, marking its status as a decacorn.
- Inference engineering has evolved into a critical discipline, distinct from model training.
- Quantizing layers can preserve benchmark quality while increasing throughput by 20%.
- New techniques like cache-aware routing and speculative decoding enhance model performance.
- The conversation expands to include local inference and the convergence of training and inference.
DeepSignal Analysis
What happened
Baseten has raised $13 billion, positioning itself as a significant player in AI infrastructure. Philip Kiely and Ali Taha discussed the emergence of inference engineering as a distinct discipline, emphasizing quantization's potential to enhance throughput by 20% without degrading model quality. Their insights reflect a shift in focus from model training to optimizing inference processes.
Key evidence
- Baseten raised $13 billion, joining the ranks of AI infrastructure decacorns alongside Nvidia and Intel.
- Philip Kiely noted that inference engineering has evolved into a critical discipline, addressing how to efficiently turn trained models into scalable products.
- In experiments with the GLM-5.2 model, quantization techniques improved throughput by 20% while maintaining benchmark quality.
Why it matters
The rise of Baseten and the focus on inference engineering highlight a significant evolution in AI infrastructure. As companies increasingly prioritize efficient model deployment, understanding inference optimization becomes crucial. The ability to enhance throughput without sacrificing quality could lead to more effective AI applications across various industries, potentially reshaping market dynamics.
Source Excerpt
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from Latent Space
See more →
Ontologies Are So Back: Why AI Agents Are Reviving the Semantic Web
Frank Coyle at AIEWF 2026 emphasized the revival of ontologies as essential 'logical guardrails' for effective AI agents, integrating them with for better reasoning. Neo4j's Emil Eifrem highlighted three ontology types to enhance agent scalability, while Kingsley Idehen discussed the challenges and benefits of maintaining ontologies in AI systems.
![[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition](https://substackcdn.com/image/fetch/$s_!8D6O!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fpbs.substack.com%2Fmedia%2FHMuQw2BXUAAJaQd.png)
