Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

NVIDIA cuts voltage droop by over 60% with Groq 3 LPX deterministic execution to get more inference tokens per megawatt

Synopsis

NVIDIA describes the deterministic execution model of the Groq 3 LPX low-latency accelerator within its Vera Rubin platform: the LPU compiler schedules compute and data movement down to the clock cycle, letting it predict per-cycle current demand and use Preemptive Power (PEP) and Clock Period Synthesis (CPS) to reduce voltage droop and shrink the voltage guardband, with internal testing showing over 60% less voltage drop, an estimated high single-digit percentage decrease in the baseline voltage that must be continually supplied, a potentially low-double-digit percentage reduction in power for the same workload, and up to 35x higher throughput per megawatt at the platform level versus the previous-generation GB200 NVL72 for 2T+ parameter models at long context and high interactivity.

AI-generated editorial illustration: How NVIDIA Groq 3 LPX Deterministic Execution Drives Power-Efficient High-Interactivity Inference on NVIDIA Vera Rubin

Interpretation

The Groq 3 LPX deterministic execution model lets the LPU compiler produce a clock-cycle-exact schedule before a workload runs, specifying when each piece of data moves to which compute unit and when each operation executes, extending across all 256 LPU chips in the rack. Unlike most chips that dynamically switch operations at runtime as resources become available, this fixes the timing of compute and data movement in advance, so individual computations, memory reads, and interchip communications take the same number of clock cycles from run to run. The article grounds this determinism in hardware: each LPU has a small set of distinct units (MXM for matrix multiplication, VXM for vector operations, SXM for transposing and reshaping), hardware that keeps these units synchronized down to the clock cycle, hierarchy-free on-chip SRAM banks, and LPUs connected directly to each other rather than through an intermediary.

Because the schedule is cycle-exact, the compiler can predict the workload's current demand over time and enable Preemptive Power (PEP), which prepares the power-delivery system before demand arrives, and Clock Period Synthesis (CPS), which shapes how abruptly demand rises and falls by lengthening or shortening individual clock cycles. Chips normally must continually provide a voltage guardband against unpredictable current spikes; here both power-delivery preparation and clock-cycle length become part of compile-time scheduling, allowing the guardband to shrink. The article gives a mechanism-level causal chain: larger current change and higher di/dt cause larger voltage droop; PEP leaves decoupling capacitors less of a gap to make up, and CPS lengthens the cycles with the very highest current spikes. CPS depends on the Groq 3 LPX plesiosynchronous clock system, which keeps clocks synchronized between chips and compute elements synchronized within chips.

Internal testing shows this approach leads to over 60% less voltage drop, and it is estimated to yield a high single-digit percentage decrease in the baseline voltage the electrical system must continually provide; because power is proportional to the square of voltage, the percentage decrease in power is even higher, all without impacting the workload. The article extends deterministic scheduling from a performance lever to a power lever, stating that Groq 3 LPX deterministic execution can reduce the power required for the same workload by a potentially low-double-digit percentage compared with a similarly specified, nondeterministic system. Evidence comes from internal testing on Groq 3 LPX systems, with no test scale, comparison configuration, or statistical detail given; the voltage-to-power conversion rests on the article's stated square relationship, where 10% more voltage continually supplied means 21% more power.

At the platform level, pairing Groq 3 LPX with Vera Rubin NVL72 enables up to 35x higher throughput per megawatt compared with the previous-generation NVIDIA GB200 NVL72 for 2T+ parameter models at long context and high interactivity; at the factory level, NVIDIA DSX MaxLPS can provision up to 40% more GPUs within the same site-power envelope and deliver 35% higher token throughput. These figures place in-rack deterministic power control alongside factory-level and rack-level power management under one goal: producing more tokens within a fixed power budget. The article presents these as platform comparisons and factory-level software effects without stating measurement conditions, model lists, or benchmark details; the determinism-enabled power controls are slated to come to the NVIDIA Vera Rubin platform in H2 2026.

Perspective

This article is aimed at readers planning AI factory power and deployment density, explaining how deterministic execution lets the compiler lay out a cycle-level schedule before a workload runs and use PEP and CPS within the rack to shrink the voltage guardband. It applies to high-interactivity, long-context inference with 2T+ parameter models, and it layers with factory-level DSX MaxLPS and rack-level Intelligent Power Smoothing. For a reader, it offers a way to turn power-supply margin into usable compute: first make timing predictable, then use that predictability for power scheduling.

The benefit figures (over 60% less voltage drop, high single-digit baseline voltage decrease, low-double-digit power reduction, 35x throughput per megawatt, 40% more GPUs and 35% higher token throughput) come from internal testing and estimates, with no stated test scale, comparison system configuration, model list, or measurement conditions, so how broadly they reproduce remains an open question. The approach also demands close compiler-hardware co-design, leaving open how similar effects could be obtained on nondeterministic architectures and how it will perform when it lands in H2 2026.

Sources