Skip to main content
Back to timeline
NVIDIA BlogSource publication:

NVIDIA frames roughly $60M-per-megawatt AI factory returns around three factors, citing SemiAnalysis data that Vera Rubin NVL72 delivers over 30x throughput per megawatt and up to 45x lower cost per million tokens than GB300 NVL72

Synopsis

This NVIDIA article argues that AI factory return on investment is set by three factors that cannot fully substitute for one another—earning capacity (tokens per second per megawatt within a fixed power envelope), useful life (how long hardware keeps earning), and demand (market appetite for those tokens)—and says NVIDIA AI factories maximize all three through full-stack codesign, generation-spanning CUDA software, and a general-purpose accelerated architecture, citing SemiAnalysis AgentX data that Vera Rubin NVL72 delivers over 30x higher throughput per megawatt than GB300 NVL72 and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model, while third-party figures such as A100 GPUs still in commercial service six years after shipping in 2020, CoreWeave extending bookings for

AI-generated editorial illustration: Productive, Durable, Fungible: How NVIDIA AI Factories Maximize Return on Investment

Interpretation

The article decomposes AI factory returns into earning capacity, useful life, and demand, and states that strength in one cannot fully offset weakness in another: high earning capacity counts for little if only part of output is sold, and high demand matters little if full-capacity production stops after a year. Relative to discussions centered on peak compute or unit cost alone, it places power-constrained token output, hardware depreciation horizon, and workload diversity inside a single return framework, and stresses that the three are interdependent—a factory that runs more kinds of workloads finds more demand and keeps earning year after year. This is the article's analytical framing rather than an experimental measurement; it grounds the stakes by noting that each megawatt factory costs roughly $60 million, which is why operators need a clear view of return before committing capital.

On earning capacity, the article says power is the binding constraint, making tokens per second per megawatt the number that governs earning capacity; it cites SemiAnalysis AgentX data showing Vera Rubin NVL72 delivers over 30x higher throughput per megawatt than GB300 NVL72 and up to 45x lower cost per million tokens on the DeepSeek V4 Pro model. It shifts the metric from absolute throughput to throughput per megawatt and cost per million tokens, which map directly onto revenue and margin inside a fixed power budget, and attributes the gains to extreme codesign across the full stack from models and workloads down through software to compute, networking, and memory. The figures come from third-party SemiAnalysis AgentX and are given as ranges (over 30x, up to 45x) rather than full measurement detail; the article does not show test configuration, model scale, or measurement method.

On useful life, the article says not every workload needs the newest system and the right fit depends on a workload's complexity and shape, so the previous generation keeps earning after the next arrives; the A100 GPU shipped in 2020 and is still in commercial service six years later, and CoreWeave recently extended bookings for units first introduced in 2020 through 2029. It turns depreciation schedules from an accounting assumption into an observable market phenomenon: Barkr puts useful life at five to six years for an eight-GPU H100 system and nine to 10 years for GB300 NVL72, Silicon Data shows a six-year-old A100 still worth a quarter of what it cost where a five-year schedule had it at zero more than a year ago, and Ornn Data finds the market paying 80% as much to rent an A100 on a five-year contract as on a one-month contract. Supported by multiple third-party sources (Barkr, Silicon Data, Ornn Data) plus operator depreciation-schedule extensions; this is market observation and resale-valuation evidence, and the article also points to a September 2026 Sprout analysis, 'The Productive Life of a Data Center GPU,' tracking how schedules have shifted across major operators.

On demand (fungibility), the article says NVIDIA AI factories run every type of AI model—open and proprietary—across language, vision, biology, physics, and robotics; every phase from data processing through pretraining, post-training, and inference; and every place from hyperscale and AI clouds to sovereign programs, enterprise data centers, and the edge; the same infrastructure also runs non-AI workloads such as data processing, scientific computing, simulation, and graphics. It grounds the fungibility argument in production cases: Lilly builds and runs protein, small-molecule, and genomics models on a 1,016-GPU on-premises cluster; Pinterest post-trains and deploys a vision language model across 14,000 GPUs spanning Blackwell, Hopper, and earlier architectures; Revolut processes data for billions of transaction records with cuDF; Runway trains a world model on Hopper and serves on Blackwell; Texas A&M University runs molecular simulation and AI drug discovery at 95-98% utilization across 26 projects and seven institutions; Dassault Systèmes powers virtual twin simulation behind aircraft certification at Wichita State and vehicle design at Lucid Motors; and Unilever builds product imagery from digital twins instead of photo shoots, cutting production costs in half. The cases are named customers and institutions with specific GPU counts and utilization figures, but they are company statements or vendor compilations, and the article provides no independent audit or comparison baseline.

Perspective

The article is aimed at operators, investors, and enterprise IT decision-makers evaluating megawatt-scale AI factory capital commitments, and applies to power-constrained data centers that need long depreciation horizons and multi-workload reuse. It supplies an evaluation vocabulary: tokens per second per megawatt for earning capacity, market residual values and lease prices for useful life, and the range of runnable workloads for demand. For a reader, this frame can be used to question suppliers—for example, asking for the measurement configuration behind throughput-per-megawatt claims, the depreciation assumptions, and the workload mix, rather than accepting peak compute or per-token price alone.

Several open questions remain for a careful reader: the article does not state the model scale, batch, or power configuration under which SemiAnalysis AgentX measured the 30x and 45x figures; the Barkr, Silicon Data, and Ornn Data valuations rest on secondary markets and lease quotes whose representativeness varies by region and time; the article mentions a September 2026 Sprout analysis tracking depreciation-schedule shifts but gives no specific values; and the customer cases' utilization and cost savings (such as 95-98% utilization and halved production costs) lack comparison baselines. In addition, because the article is published by a hardware vendor, readers adopting its return framework would do well to cross-check it against independent modeling and their own workload mix.

Sources