NEWS · CODING · #1378
NVIDIA publishes DIN Deploy C++ samples using ONNX Runtime + TensorRT RTX
Do Inference Now (DIN) Deploy is an open-source repository of C++ samples that pair ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux. Samples demonstrate ASR (OpenAI Whisper, Parakeet TDT, Nemotron), Meta SAM 2.1 segmentation, and FLUX.2-klein-4B image generation, include CMake presets (x86-64, Arm64), show graphics interop using ONNX Runtime 1.25 with Vulkan/DirectX, and report GPU speedups (e.g., 206x for Parakeet TDT, 39x for Nemotron streaming, SAM 2.1 at 38.3 FPS vs 0.5 FPS on CPU).
KEY POINTS
- Do Inference Now (DIN) Deploy is an open-source repository of C++ samples that pair ONNX Runtime with the NVIDIA TensorRT RTX execution provider to accelerate local AI inference on Windows and Linux.
- Samples demonstrate ASR (OpenAI Whisper, Parakeet TDT, Nemotron), Meta SAM 2.1 segmentation, and FLUX.2-klein-4B image generation, include CMake presets (x86-64, Arm64), show graphics interop using ONNX Runtime 1.25 with Vulkan/DirectX, and report GPU speedups (e.g., 206x for Parakeet TDT, 39x for Nemotron streaming, SAM 2.1 at 38.3 FPS vs 0.5 FPS on CPU).
- These samples give developers a practical, end-to-end path from Hugging Face checkpoints to native, GPU‑accelerated C++ applications using ONNX Runtime and TensorRT RTX, reducing integration friction for local AI.
WHY IT MATTERS
These samples give developers a practical, end-to-end path from Hugging Face checkpoints to native, GPU‑accelerated C++ applications using ONNX Runtime and TensorRT RTX, reducing integration friction for local AI.