This is not a consumer app, and there's no button that makes your PC smarter. The GitHub repository calls DIN Deploy "a collection of samples." It's sample code that shows how to get from a Hugging Face checkpoint to a native, GPU-accelerated executable.
How the samples are split
Every DIN Deploy sample has two halves:
- A Python exporter downloads a model checkpoint from Hugging Face and converts it into a directory of ONNX files.
- A native C++ command-line tool loads that directory with ONNX Runtime. It runs on either the TensorRT RTX execution provider (EP) or the CPU EP.
Model conversion is a one-time job for developers. Inference is what ships to users. Keeping them apart means a Windows app can consume an exported model without carrying a runtime built for that particular model.
NVIDIA says most of the sample code uses ONNX Runtime's standard session and tensor APIs. Vendor-specific code, including CUDA APIs and custom kernels, appears only in optional accelerated paths. The ORT copy-tensor API moves data between devices so the shared code doesn't have to call vendor APIs directly. In theory, any execution provider that supports the required ORT tensor APIs can run the shared code.
That doesn't mean every provider and GPU are interchangeable. NVIDIA's own TensorRT for RTX architecture documentation warns that ONNX conversion is all-or-nothing: every operation in the model must be supported by TensorRT-RTX. The code may be portable, but each provider still has to support the model's operators.
Section summary: DIN Deploy keeps export (Python) separate from inference (C++/ORT). Most code is vendor-neutral, and CUDA appears only on optional fast paths.
What the samples cover
The repository lists three groups of workloads:
| Category | Models | Notes from NVIDIA |
|---|---|---|
| Speech recognition | OpenAI Whisper (tiny, base, small, medium, large-v3, large-v3-turbo); NVIDIA Parakeet TDT 0.6B v3; NVIDIA Nemotron 3.5 ASR Streaming 0.6B | Whisper handles offline transcription. Parakeet and Nemotron provide streaming pipelines. |
| Computer vision | Meta SAM 2.1 (hiera-tiny, small, base-plus, large) | Interactive masking for images and video. The masks can be used for selection and tracking. |
| Image generation | FLUX.2-klein-4B (standard, FP8, NVFP4 variants) | Prompt-driven generation, plus Vulkan/DirectX interop and quantization with NVIDIA Model Optimizer |
The repository also includes din_base_onnx, a ResNet-18 example for developers who want to see a basic native ONNX integration before moving on to the larger models.
The benchmarks were run on DGX Spark
NVIDIA published GPU-versus-CPU figures. Audio results are expressed as multiples of real time, where higher is faster:
| Model | GPU (TensorRT RTX EP) | CPU EP |
|---|---|---|
| openai/whisper-large-v3-turbo | 58.5× real time | 3.8× |
| nvidia/nemotron-3.5-asr-streaming-0.6b | 39.01× | 3.24× |
| nvidia/parakeet-tdt-0.6b-v3 | 206.41× | 14.44× |
| facebook/sam2.1-hiera-base-plus | 38.3 FPS | 0.5 FPS |
All of these numbers come from NVIDIA's DGX Spark, a compact Arm-based AI workstation. They weren't measured on a typical Windows gaming PC or office laptop. NVIDIA doesn't publish results for consumer GeForce cards, other CPUs or other driver versions. Read the table as a sign of how big the GPU-versus-CPU gap is for these workloads, not as a prediction for your hardware.
The same caution applies to quantization. NVIDIA's chart shows FLUX.2-klein-4B running about 1.33× the BF16 speed at FP8 and 1.63× at NVFP4. The repository says these figures were measured on DGX Spark using the fully CUDA-backed pipeline. Neither the blog nor the README says that image quality stays the same at lower precision.
Quantized models drop in without code changes
The FLUX.2 sample shows post-training quantization (PTQ) with NVIDIA Model Optimizer producing a quantized ONNX model. The ONNX inputs and outputs don't change, so NVIDIA says the quantized model can replace the original without any changes to application code. NVIDIA also notes that quantization depends on the hardware. FP8 and NVFP4 gains depend on recent tensor-core hardware, so older GPUs won't see the same benefit.
This lines up with NVIDIA's TensorRT for RTX documentation, which says the runtime works with quantized ONNX models exported by TensorRT Model Optimizer or any other 3rd party quantization library.
DirectX interop is Windows-only
For pre- and post-processing around inference, the FLUX.2 sample uses ONNX Runtime's graphics interop feature, which arrived in ORT 1.25. It works with both Vulkan and DirectX. NVIDIA notes that DirectX is available only on Windows. Vulkan is the cross-platform option.
ONNX Runtime's API reference describes OrtGraphicsInteropConfig and InitGraphicsInteropForEpDevice as available since version 1.25, supporting D3D12 and Vulkan:
- On D3D12, developers can pass an
ID3D12CommandQueue*, which the provider may use for GPU-side synchronization with inference streams. - On Vulkan, the command-queue field is set to null.
- The command queue is optional. Without one, interop still works and streams use the default context.
This lets an app keep textures and buffers on the GPU instead of copying them back to system memory between rendering and inference. It's an explicit API with synchronization you have to set up yourself. Your renderer won't start sharing resources with the model on its own.
Where Windows ML fits
NVIDIA mentions that the same ONNX Runtime API is available through WinML 2.0, and Microsoft's own servicing records show how closely the two are tied. Microsoft describes the NVIDIA TensorRT RTX Execution Provider as an ONNX Runtime/Windows machine learning execution provider designed specifically to accelerate ONNX model inference on NVIDIA RTX GPUs for client-centric, end-user PC scenarios. Microsoft ships this component through Windows Update as a separate package. KB5121778 moved it to version 2.2607.0.0 and includes improvements to the execution provider component for Windows 11, version 26H1. Earlier, KB5089168 delivered version 2.2604.1.0 to Windows 11 24H2 and 25H2.
NVIDIA's documentation states that TensorRT-RTX is the default EP when Windows ML runs on supported RTX hardware, and recommends the direct ORT execution-provider route when your application already integrates ONNX Runtime.
That leaves two deployment options:
- DIN Deploy's default path: the app bundles a specific ONNX Runtime build and the TensorRT RTX EP. You control the exact versions.
- The Windows ML path: Windows supplies and services the execution provider, so your app can get improvements through Windows Update.
DIN Deploy doesn't include a WinML-specific build. Moving from option 1 to option 2 is something you'd test yourself, and NVIDIA doesn't claim it's a guaranteed straight swap.
One more detail matters for packaging. ONNX Runtime's TensorRT RTX documentation says the built-in TensorRT RTX Execution Provider in the ONNX Runtime repository is deprecated. We strongly recommend using the standalone EP ABI plugin instead. DIN Deploy's README says CMake builds the TensorRT RTX EP from source by default. Check which EP form the build produces before you decide how to ship it.
Section summary: DIN Deploy shows the "bring your own ORT" route. Windows ML offers an OS-serviced alternative built on the same API, but you'll need to test before treating them as equivalent.
Building the samples
NVIDIA's blog makes setup sound like one CMake command. The README lists more prerequisites than that:
Requirements
- A C++20 compiler and CMake 3.24 or newer
- Python 3.12 or newer for the model exporters
- The CUDA Toolkit and TensorRT RTX SDK for TensorRT RTX EP builds
- A preset from
CMakePresets.json, such aswindows-x64,linux-x64or an Arm64 variant
Steps
- Create a Python environment and install the repository with the extra for your model family. For example:
python -m pip install -e ".[parakeet]". The extras are.[nemotron]and.[flux]. - Export the model to a directory using the command in that model's README.
- Configure:
cmake --preset <preset>. On the first configure, supported presets download TensorRT RTX automatically. CMake also downloads ONNX Runtime, defaulting to 1.27.0 unless you setONNXRUNTIME_ROOT. - Build:
cmake --build out/build/<preset> --config Release. To build a single sample, add a target such as--target din_asr_whisper_cli,din_asr_parakeet_tdt_cli,din_asr_nemotron_cliordin_flux2_cli. - Run the
din_*_cliexecutable and pass your exported folder with--model-dir. Each CLI lists its options with--help.
What success looks like: with multi-config generators such as Visual Studio, the executables and required runtime libraries end up in out/build/<preset>/bin/<configuration>/.
Useful CMake switches
DIN_BUILD_TRT_RTX_EP(default ON): set it to OFF if you've downloaded a compatible provider library and placed it in the executable directory.DIN_ENABLE_NVTX(default ON): set it to OFF to build without NVTX profiling instrumentation.TRT_RTX_ROOT: points the build at a local TensorRT RTX SDK instead of the download. The folder must containinclude/andlib/.
The ORT 1.27.0 default refers to the version CMake downloads. The 1.25 figure is when graphics interop first appeared in ONNX Runtime. The two numbers describe different things.
Licensing
DIN Deploy is licensed under Apache 2.0, and third-party software and model licenses are listed separately in THIRD_PARTY_NOTICES.md. Whisper, SAM 2.1, FLUX.2 and NVIDIA's own speech models each have their own terms. Check those before you ship an exported model in a commercial product.
The takeaway
NVIDIA is pitching DIN Deploy as a cross-vendor framework, but its fastest paths run on NVIDIA hardware. The core design still holds up: export to ONNX, keep the C++ code on standard ORT APIs, and limit CUDA to optional extras. The Windows-specific parts are DirectX 12 interop and the WinML 2.0 compatibility. With Microsoft already servicing the TensorRT RTX provider through Windows Update, these samples are a reasonable starting point for C++ developers adding local transcription, segmentation or image generation to Windows apps. Run your own benchmarks on your users' hardware rather than relying on the DGX Spark figures.
References
- Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples | NVIDIA Technical Blog - NVIDIA Developer NVIDIA Developer · 2026-10-01T17:59:27+00:00
- DIN Deploy — DO INFERENCE NOW github.com
- KB5121778: Nvidia TensorRT-RTX Execution Provider update (version 2.2607.0.0) support.microsoft.com