Skip to main content
Back to timeline
NVIDIA 开发者技术博客Source publication:

NVIDIA and Nscale test DSX MaxLPS in Iceland: GPUs rise from 140 to 192 and aggregate throughput gains 49.2% under the same 264.4 kW budget

Synopsis

NVIDIA and Nscale jointly evaluated DSX MaxLPS policy-governed dynamic power allocation with Kimi K2.5 (FP4) inference workloads on NVIDIA GB300 NVL72 systems at Nscale's data center at the Verne campus in Keflavík, Iceland: under the same 264.4 kW approved power budget, managed GPUs rose from 140 to 192 (+37.1%), aggregate throughput rose from 1,084,503 to 1,618,443 tokens/s (+49.2%), throughput per provisioned watt rose from 4.10 to 6.12 tokens/s/W (+49.2%), median and P75 latency stayed within 5% of baseline, while P99 time to first token increased 17% from the 15.7-second baseline.

AI-generated editorial illustration: How NVIDIA DSX MaxLPS Maximizes AI Factory Throughput and Efficiency

Interpretation

Under a fixed power budget, dynamic power allocation turns idle headroom into usable compute: managed GPUs rose from 140 to 192, aggregate throughput rose 49.2%, and throughput per provisioned watt rose from 4.10 to 6.12 tokens/s/W. Relative to static per-node peak reservation, this evaluation provides a measured comparison of capacity and throughput under the same 264.4 kW budget rather than only describing the mechanism. The joint evaluation collected GPU, node, rack, and group power telemetry at Nscale's data center and measured normalized aggregate and per-instance throughput, time to first token, end-to-end latency, and interactivity; both configurations used the same 264.4 kW provisioned budget as a fixed denominator.

Scaling did not sacrifice service levels of existing instances: high-throughput output per instance moved from 59,153 to 59,220 tokens/s (+0.1%), and the low-latency instance held at 2,265 tokens/s (0%). This shows that adding a third 52-GPU high-throughput instance left the throughput of the existing high-throughput and low-latency instances effectively unchanged at the displayed precision. Per-instance throughput was measured and compared across both configurations, and distributed workloads were confined to a single rack in both configurations to control for cross-rack performance differences.

The gain comes with observable costs: mean GPU power rose from 97.0 kW to 131.8 kW (+35.9%), total measured power rose from 166.2 kW to 198.9 kW (+19.7%), and power-budget utilization rose from 62.9% to 75.2%. This frames dynamic allocation as making existing engineering trade-offs observable and controllable at the fleet level, not as eliminating them. Power figures come from GPU, CPU, and rack-level telemetry, with site-level telemetry used to verify rack-level power measurements.

Tail latency must be validated separately: median and P75 latency remained within 5% of baseline, but P99 time to first token increased 17%. This signals that stable median latency can mask changes in tail latency, and that production acceptance criteria should be defined as part of the study. The evaluation reports both P75 and P99 percentile results and gives the magnitude of the P99 time-to-first-token change relative to the 15.7-second baseline.

Perspective

This evaluation is aimed at operators and infrastructure teams running AI factories in power-constrained settings, where the managed boundary is defined, telemetry is reliable, and an aggregate power budget is the constraint. It offers a repeatable staged validation path: define the managed boundary, establish a representative baseline, introduce policies conservatively, add capacity and test each stage, and set production operating limits. It also notes that dynamic allocation is an operational capability while the site must still size electrical distribution, cooling, network fabric, floor space, and rack positions for the validated lifecycle target; for future NVIDIA Vera Rubin NVL72, any capacity projection should remain separate from this measured GB300 NVL72 evaluation.

Readers should still watch: available headroom depends on the workload mix, since complementary power profiles create more opportunity than workloads that peak simultaneously; dynamic allocation depends on reliable telemetry, and missing, delayed, or incorrectly mapped measurements can undermine fleet-level decisions; stable median latency can mask changes in tail latency, as seen in the 17% increase in P99 time to first token. In addition, this evaluation ran across four racks with each distributed workload confined to a single rack, so equivalent performance across racks needs separate validation, and production acceptance criteria should be defined as part of the study.

Sources