Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
Quick Answer
Netflix has detailed its in-house LLM serving platform, integrating Triton and vLLM to manage model inference across CPUs and GPUs.
Quick Take
The architecture supports real-time and batch workloads while addressing compatibility issues and custom model integration, ensuring a stable deployment environment despite evolving technologies.
Key Points
- Netflix's platform utilizes Triton for model management and vLLM for inference execution.
- Compatibility issues between Triton and vLLM versions can hinder deployment success.
- Constrained decoding ensures model outputs meet specific formats like valid JSON.
- Versioned deployments allow for seamless migration between model revisions.
- The architecture separates application integrations from underlying model and runtime complexities.
DeepSignal Analysis
What happened
Netflix has outlined its internal LLM serving platform, which integrates Triton and vLLM for managing model inference across CPUs and GPUs. The architecture supports both real-time and batch workloads while addressing compatibility issues and custom model integration, ensuring a stable deployment environment.
Key evidence
- Netflix's platform allows smaller models to run in-process on CPUs, while larger requests are managed by Triton, which handles model loading and GPU scheduling.
- The company found that mismatched versions of Triton and vLLM can prevent deployments from loading, necessitating compatibility testing.
- Netflix's deployment strategies include Red-Black and Versioned deployments, which allow for separate availability of old and new model revisions.
Why it matters
The development of Netflix's LLM serving platform highlights the complexities of integrating various model sizes and hardware requirements. By addressing these challenges, Netflix aims to maintain a consistent production workflow, which is crucial for delivering reliable AI services. This approach reflects broader industry trends towards modular architectures that can adapt to evolving technologies without disrupting service.
📖 Reader Mode
~3 min readNetflix has described the production lessons behind bringing LLM inference into its internal serving platform, including the challenges of supporting different model sizes, hardware requirements, and rapidly evolving inference engines.
The company’s account covers the architectural choices and operational work involved in running real-time and batch workloads across CPUs and GPUs, from model packaging and deployment to constrained decoding and version compatibility.
The platform builds on Netflix’s existing JVM-based serving layer, which continues to handle routing, feature retrieval, candidate generation, post-processing, and logging. Smaller models can run in-process on CPUs, while larger requests are delegated to MSS, where Triton takes over model loading, batching, GPU scheduling, and multi-framework serving. This allows the surrounding production workflow to remain consistent even as inference moves between local and remote hardware.
Within the GPU path, Netflix selected vLLM for its operational fit and extensibility while retaining Triton’s model-management and scheduling responsibilities. Triton controls the serving environment around the model, whereas vLLM performs inference and provides extension mechanisms for custom behaviour. Netflix reports that mismatched Triton and vLLM versions can prevent deployments from loading, requiring compatible releases to be tested and pinned together.

Custom models introduced another integration challenge. Hugging Face compatibility in vLLM was insufficient for some Netflix models, so the company used vLLM extension points for custom architectures and decoding behaviour.
Netflix also compared two Triton packaging approaches: Triton’s Python backend and its vLLM backend. The company reports that the vLLM-backend approach allows models and frontends to evolve more independently than the Python-backend option. The choice affects how tightly a model is coupled to its serving environment rather than which engine performs inference.
The common serving interface did not eliminate differences between the underlying engines. Although Triton exposed an OpenAI-compatible API alongside KServe HTTP and gRPC frontends, Netflix still encountered gaps in how some features were handled across those integrations.
Constrained decoding was one example. It allows Netflix to force model responses into formats such as valid JSON by filtering the tokens the model may generate at each step. Because those rules depend on everything generated so far, the decoder must maintain state throughout the request. When vLLM pauses and later resumes a request to manage GPU resources, that state can fall out of sync with the token history, so
Netflix added logic to detect the change and rebuild it before generation continues.
Compatibility also affected deployment. Netflix pins tested Triton and vLLM versions together to prevent backend-loading failures, while Red-Black and Versioned deployment strategies handle changes at the model level. Versioned deployments keep old and new revisions available separately, allowing consumers to migrate after adapting to incompatible input or output schemas.
Uber has described a related approach at the application boundary. Its generative AI gateway presents an OpenAI-compatible interface across externally hosted and internally managed models, while centralising concerns including authentication, caching, observability, and routing. The implementation differs from Netflix’s serving platform, but both separate application integrations from the models, runtimes, and hosting environments behind them.
Netflix’s experience shows how a common serving interface can sit above several distinct layers. This type of architecture reflects a broader effort to give application teams a stable integration surface while model providers and serving runtimes continue to change. Netflix’s account also shows that the abstraction does not remove the underlying work: packaging, compatibility controls, constrained decoding, and deployment isolation still require engineering at each layer.
About the Author
Matt Foster
Show moreShow less
— Originally published at infoq.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from InfoQ AI, ML & Data Engineering
See more →Google Cloud Workbench Notebooks Extension Connects VS Code to Google Cloud's Jupyter Notebooks
The Google Cloud Workbench Notebooks extension for VS Code allows developers to seamlessly connect their local IDE to managed Jupyter notebook environments on Google Cloud, enhancing ML workflow efficiency. This integration eliminates context switching, enabling smooth transitions from local experimentation to high-performance cloud computing.

