NVIDIA brings TensorRT multi-device inference into Dynamo-Triton 26.07 behind a single gRPC endpoint, cutting Cosmos 3 Nano video generation end-to-end latency from 156.595 s to 34.183 s
Synopsis
NVIDIA describes integrating TensorRT multi-device inference (fully supported starting with TensorRT 11.0) into Dynamo-Triton 26.07, where one KIND_MODEL instance can own multiple GPUs, create per-rank TensorRT execution contexts, CUDA streams, and NCCL communicators, and launch the ranks together, so an application calls one named model through one gRPC endpoint; in a Cosmos 3 Nano video generation example, a 36-layer denoising transformer uses Ulysses context parallelism to distribute 44,160 video tokens across as many as eight GPUs, dropping end-to-end latency from 156.595 s on one GPU to 34.183 s on eight (4.58x), with 6.09x transformer RPC speedup and CP2/CP4/CP8 all passing the configured MAE ≤ 25 and PSNR ≥ 18 dB thresholds.
Interpretation
Dynamo-Triton 26.07 enables the multi-device inference capability of the TensorRT backend: one Triton KIND_MODEL instance can own multiple GPUs, create per-rank TensorRT execution contexts, CUDA streams, and NCCL communicators, and launch the ranks together for each request, while the application calls one named model through a gRPC endpoint instead of coordinating GPU ranks itself. Previously there was a gap between multi-GPU acceleration and a consumable inference service; this integration moves rank and communicator lifecycle code out of the client and lets the engine be packaged as a versioned Triton model while keeping the application interface and surrounding workflow stable. The text includes an excerpt of the generated CP8 config.pbtxt showing backend tensorrt, instance_group KIND_MODEL count 1, enable_multi_device set to true, and multi_device_gpus set to 0 through 7; this is configuration-level evidence rather than an independent third-party evaluation.
In the Cosmos 3 Nano video generation example, end-to-end latency falls as GPUs increase: 156.595 s on one GPU, 87.999 s at CP2, 53.093 s at CP4, and 34.183 s at CP8, corresponding to 1.00x, 1.78x, 2.95x, and 4.58x end-to-end speedup; mean transformer RPC time drops from 146.192 s to 23.993 s, a 6.09x RPC speedup. The example quantifies multi-device inference on a long-sequence generation workload: transformer RPCs account for 93.4% of generation time on one GPU and fall to 70.2% at CP8, showing the gain concentrates in the distributed transformer stage. All four variants ran on the same healthy eight-GPU NVIDIA system with 1280×720 output, 189 frames at 24 FPS, and 35 denoising steps; each result includes one warm-up followed by five measured complete generations, and timing covers prompt work, the 70 Dynamo-Triton calls, CFG and scheduler updates, VAE decode, and frame postprocessing, while excluding model loading and mp4 encoding.
The distributed Ulysses graph is compiled into each TensorRT plan before deployment: the engine is exported from PyTorch and compiled with Torch-TensorRT, three local converters lower export-carrier operations to the TensorRT public distributed-collective layer (reduce-scatter, all-to-all, all-gather), and each accepted context-parallel plan contains two initial reduce-scatters, three all-to-alls in each of 36 transformer layers, and one final all-gather, for a topology of two reduce-scatters plus 108 all-to-alls plus one all-gather. The text states explicitly that Dynamo-Triton configuration activates that plan and does not convert a single-device engine into a distributed engine, so the distributed topology is fixed at compile time. This is a mechanistic description of compilation and communication topology, paired with the CP8 configuration excerpt; it is implementation-level explanation rather than a comparison against other frameworks.
Context-parallel outputs match the single-device result within configured thresholds: every variant used the same seed and generation profile, validation sampled frames 0, 47, 94, 141, and 188, checked format and temporal variation, and compared each context-parallel output with the single-device result; CP2 and CP4 measured MAE 12.759 and PSNR 21.111 dB, and CP8 measured MAE 16.316 and PSNR 19.400 dB, all passing the MAE ≤ 25 and PSNR ≥ 18 dB thresholds. The text does not claim pixel-identical outputs, using numeric thresholds and a contact sheet showing the same coherent action across the clip (a robot arm cleaning a plate) as quality evidence. The evidence is a numeric comparison over five sampled frames plus a qualitative observation, with a limited number of sample points and thresholds set by configuration.
Perspective
The capability targets teams that need shorter per-request latency and are willing to trade additional GPU resources for response time, fitting latency-sensitive generative media workflows such as reducing user wait time and accelerating review-and-refine cycles. The results apply to the Cosmos 3 Nano long-sequence workload described and to the same eight-GPU system, with the distributed graph compiled into a versioned TensorRT plan that Dynamo-Triton activates through a KIND_MODEL instance; on the application side only one gRPC endpoint is called. The text gives a reproduction path: download NVIDIA Dynamo-Triton 26.07 from NGC and use the linked TensorRT, Torch-TensorRT, Diffusers, and Cosmos resources.
The text states that this benchmark does not measure concurrent request throughput, cost per generated video, or total cost of ownership, so teams still need to evaluate the resource-for-latency trade-off against their own SLOs and deployment economics. Timing excludes model loading and mp4 encoding, so the end-to-end numbers correspond to the defined measurement scope. Outputs are not claimed to be pixel-identical, and quality evidence comes from MAE and PSNR over five sampled frames plus a qualitative contact-sheet observation. In addition, the results come from a single model example on one eight-GPU system, so extending them to other workloads, hardware scales, or parallelism strategies still requires independent verification.
