Vibeleaderboard
← All Intel
Intel / article

Build Local AI Apps with C++ and NVIDIA TensorRT RTX Samples

Source
Luca Spindler
Author
Luca Spindler
Date
Key takeaways · AI-distilled
  • Each DIN Deploy sample is split in two: a Python exporter converts a Hugging Face checkpoint to ONNX, and a native C++ CLI built on ONNX Runtime runs it, so the app needs no model-specific runtime.
  • Shared sample code sticks to ONNX Runtime session and tensor APIs. CUDA APIs and kernels appear only in optional accelerated paths, so any execution provider that supports those tensor APIs can run the common code.
  • NVIDIA's own DGX Spark numbers put Whisper large-v3-turbo at 58.5x real time on GPU vs 3.8x on CPU, and Parakeet TDT 0.6B at about 206x vs 14x.
  • Quantizing FLUX.2-klein with NVIDIA Model Optimizer keeps the ONNX interfaces unchanged, so the model is a drop-in swap with no application-code changes, though NVIDIA notes quantization is hardware-dependent.
  • The FLUX.2 sample uses ONNX Runtime 1.25 graphics interop with Vulkan and DirectX (DirectX on Windows only) so pre- and postprocessing can work on GPU-resident resources.
Terms in this piece · Glossary
  • inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
Why it matters

Shows how to ship local speech, segmentation and image-generation models from native C++ with GPU acceleration, with measured speedups (SAM 2.1 at 38.3 FPS on GPU vs 0.5 on CPU).

Read the source developer.nvidia.com
Recommended reads
Comments

Checking sign-in…

Loading comments…