Skip to main content
Back to timeline
NVIDIA ResearchSource publication:

NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Synopsis

NVIDIA's first MLPerf Inference v6.1 preview submission of Vera Rubin NVL72 reports up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1, while GB300 NVL72 reached 99% scaling efficiency across 288 GPUs in four racks and software optimizations delivered up to 1.6x over v6.0.

AI-generated editorial illustration: NVIDIA Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut

Interpretation

Vera Rubin NVL72's first MLPerf Inference preview submission leads on two demanding benchmarks: up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios using vLLM with NVIDIA Dynamo, and up to 2.5x on DeepSeek-R1 using TensorRT-LLM. This is the platform's first MLPerf Inference preview submission, so it provides the first comparable throughput reference for Vera Rubin NVL72 against the prior-generation GB300 NVL72. Drawn from MLPerf Inference v6.1 Closed Division submissions (entries 6.1-0106 and 6.1-0074), retrieved from mlcommons.org on Sep 16, 2026; the text frames these as preview results and notes performance will improve with continued software optimization.

GB300 NVL72 scaled from a single rack of 72 GPUs to four racks of 288 GPUs on DeepSeek-R1, achieving 99% scaling efficiency in the offline scenario, with throughput growing nearly in proportion to the hardware added. It answers the infrastructure-productivity question of whether added GPUs translate proportionally into throughput, using a single-rack-to-four-rack comparison on the same benchmark rather than a single peak number. MLPerf Inference v6.1 Closed Division submissions (entries 6.1-0073 and 6.1-0074); the text attributes the efficiency to high-bandwidth low-latency scale-up interconnects, high-bandwidth inter-rack networking and cross-node request orchestration.

Software optimization drives continuous gains: in v6.1, GB300 NVL72 performance on Qwen3-VL improved up to 1.6x over v6.0 through lower KV cache precision, additional kernel fusion, better kernels and disaggregated serving with vLLM and NVIDIA Dynamo; post-submission work on GPT-OSS-120B and DLRMv3 shows further gains. It attributes performance gains to reusable software-stack improvements (precision, kernels, serving architecture), showing that the same hardware keeps yielding returns across versions rather than relying only on new silicon. The 1.6x v6.1-over-v6.0 gain is a submitted result; the post-deadline GPT-OSS-120B and DLRMv3 results are explicitly labeled as not yet verified by MLCommons.

Measurement for agentic inference is taking shape: in preview testing on SemiAnalysis AgentX, Vera Rubin NVL72 delivered 30x better performance than GB300 NVL72, and the upcoming MLPerf Endpoints benchmark will bring standardized measurement to agentic inference workloads. It extends the evaluation target from traditional throughput to agentic workloads that reason, plan and act across multiple steps, while noting that standardized benchmarks are still being established. The AgentX figure comes from preview testing rather than an MLPerf submission, and MLPerf Endpoints is described as upcoming, so this part is a directional signal rather than an established benchmark result.

Perspective

This article serves readers evaluating rack-scale inference infrastructure: it provides throughput multiples of Vera Rubin NVL72 over GB300 NVL72, GB300 NVL72's scaling efficiency from 72 to 288 GPUs, and v6.1-over-v6.0 software gains, usable for initial capacity and cost-per-token estimates. Its setting is the specific benchmarks of MLPerf Inference v6.1 Closed Division and NVIDIA's own hardware-software stack (vLLM, Dynamo, TensorRT-LLM, NVFP4, NVLink), spanning deployments from Jetson AGX Thor edge devices to NVL72 racks.

Readers should still watch that Vera Rubin NVL72 results are described as preview and final submitted scores may differ; that post-deadline results on GPT-OSS-120B and DLRMv3 are not yet verified by MLCommons; that the 30x AgentX figure comes from preview testing rather than an MLPerf submission; and that the MLPerf Endpoints benchmark is not yet released, leaving standardized measurement of agentic inference still to be defined. In addition, this is a news-style article that does not provide absolute throughput values, power or cost data per benchmark, so cost per token cannot be independently reconstructed from it.

Sources