
Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples
Quick Answer
NVIDIA's DIN Deploy offers C++ samples for integrating AI models with ONNX Runtime and TensorRT RTX, enabling efficient local applications on Windows and Linux.
Quick Take
Performance benchmarks show significant GPU acceleration, with models like OpenAI's Whisper achieving 58.5x speedup over CPU. The toolkit supports various AI tasks, including ASR and image segmentation.
Key Points
- DIN Deploy bridges AI model integration with C++ using ONNX Runtime and TensorRT RTX.
- Supports automatic speech recognition with models like OpenAI Whisper and NVIDIA Parakeet.
- Performance benchmarks show GPUs outperform CPUs significantly, e.g., Whisper at 58.5x speedup.
- FLUX.2 sample demonstrates prompt-driven image generation with Vulkan and DirectX integration.
- CMake presets simplify setup for Windows, Linux, and Arm64 environments.
📖 Reader Mode
~3 min readAdding AI models to local applications requires a portable model format, a reliable runtime, and acceleration that works across target systems.
Do Inference Now (DIN) Deploy is an open-source collection of practical C++ samples that bridges that gap. It combines ONNX Runtime with the NVIDIA TensorRT RTX execution provider to help developers move from a model checkpoint to a native, hardware-accelerated application on Windows and Linux. The same ONNX Runtime API can also be accessed through WinML 2.0.
From ONNX export to C++ implementation
Each DIN Deploy sample starts with a Python exporter that downloads a model checkpoint from Hugging Face and converts it into an ONNX artifact. The application side is a native C++ CLI built on ONNX Runtime (ORT). That split keeps model conversion separate from deployment logic, and developers can take an exported model into a local application without requiring a model-specific runtime.
Most sample code uses ONNX Runtime session and tensor APIs in C++. Vendor-specific code, including CUDA APIs and kernels, appears only in optional accelerated paths. Execution providers that support the required ONNX Runtime tensor APIs can run the shared code.
ORT’s copy tensor API keeps data locality manageable without dedicated vendor API usage in shared code.
For preprocessing and postprocessing around exported model inference, the FLUX.2 sample uses ONNX Runtime’s graphics interop capability, introduced in version 1.25, with Vulkan and DirectX for sampling.
The repository provides CMake presets for Windows and Linux, including Arm64 variants. DirectX is available only on Windows.
AI tasks supported by DIN Deploy
For automatic speech recognition (ASR), DIN Deploy supports offline and streaming pipelines. OpenAI Whisper covers offline transcription across model sizes, while NVIDIA Parakeet TDT and NVIDIA Nemotron ASR Streaming provide streaming pipelines.
The samples show how to move audio through a native application and return transcription results while using GPU acceleration where it is available.
Meta SAM 2.1 samples support interactive masking for images and video. They turn model outputs into segmentation masks that native applications can use for selection, tracking, and other computer-vision workflows.
Table 1 compares GPU and CPU performance for selected DIN Deploy workloads measured on DGX Spark.
| Model | GPU (DGX Spark) | CPU (DGX Spark) |
|---|---|---|
openai/whisper-large-v3-turbo | 58.5x | 3.8x |
nvidia/nemotron-3.5-asr-streaming-0.6b | 39.01x | 3.24x |
nvidia/parakeet-tdt-0.6b-v3 | 206.41x | 14.44x |
facebook/sam2.1-hiera-base-plus | 38.3 FPS | 0.5 FPS |
The FLUX.2-klein-4B sample provides prompt-driven image generation. The sample includes graphics-API interops with Vulkan and DirectX, allowing applications to integrate GPU-resident resources with a cross-vendor shader interface. It also shows how post-training quantization (PTQ) with NVIDIA Model Optimizer produces a quantized ONNX model. Quantization is hardware-dependent, but due to ONNX interfaces remaining unchanged, the quantized model is a drop-in replacement requiring no application-code changes.

Get started with DIN Deploy
Start with the repository’s CMake presets for Windows, Linux, x86-64, and Arm64. After configuring and building the project, export a model to ONNX and run the CLI with TensorRT RTX, or copy the code into your own application and use the pipeline implementations. CMake downloads ONNX Runtime and TensorRT RTX by default.
Learn more about TensorRT for RTX, NVIDIA Local AI, and the DIN Deploy repository, and the NVIDIA blog series on model quantization.
About the Authors
— Originally published at developer.nvidia.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from NVIDIA Developer Blog
See more →
Synthetic Data Generation for Financial AI Research with NVIDIA NeMo
NVIDIA's NeMo pipeline generates 502,536 unique financial news headlines in 82 iterations, addressing data imbalance in financial NLP. The iterative approach uses semantic deduplication and category-weighted sampling to enhance diversity and relevance in generated content.

