CoreWeave Puts NVIDIA Vera Rubin NVL72 Into Production as Cognition Measures Up to 4.8x Higher Total Token Throughput on SWE-2 Inference
Synopsis
At Fully Connected, CoreWeave announced availability of NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet on CoreWeave Cloud, with first customer Cognition benchmarking a real-world software engineering workload built from a sampled subset of FrontierCode tasks and seeing up to a 4.8x increase in total token throughput for SWE-2 inference over GB200 NVL72; CoreWeave also said it will offer NVIDIA Vera, described as the first CPU built for AI agents, and launched CoreWeave Forge, a connected environment unifying Weights & Biases, OpenPipe post-training expertise and the open source marimo notebook project.
Interpretation
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet networking on CoreWeave Cloud, describing itself as one of the first cloud providers to deliver the platform into customers' hands. The platform had not previously been delivered in production form on a cloud; the text says CoreWeave stood up a production Vera Rubin cluster for Cognition in days after receiving its first production racks. Based on the announcement and deployment narrative plus a quote from Ian Buck, NVIDIA vice president of hyperscale and high-performance computing; this is vendor release information without third-party reproduction.
Cognition benchmarked Vera Rubin NVL72 against a GB200 NVL72 baseline on a real-world software engineering workload and, in early tests, saw up to a 4.8x increase in total token throughput for SWE-2 inference workloads. The comparison used a sampled subset of FrontierCode tasks solved by deployed AI agents rather than a synthetic benchmark; the text says Cognition scaled to thousands of GPUs on CoreWeave in nine months. The text explicitly labels these as early tests and reports only a relative multiple, with no absolute throughput, task counts or statistical uncertainty.
CoreWeave will offer NVIDIA Vera CPU, with 128 CPUs and 11,264 cores in a single rack, enough for more than 11,000 concurrent environments at one core each; testing showed more than 3x faster agent sandbox startup times and a 1.7x performance gain on Terminal-Bench across all passing tasks. It pairs an agent-oriented CPU with hardware-isolated CoreWeave Sandboxes so the thousands of isolated environments needed for post-training can run alongside the training jobs they support. Figures come from CoreWeave's own testing; the text does not state test scale, repetitions or baseline configuration.
CoreWeave Forge unifies Weights & Biases, post-training expertise from OpenPipe and the open source marimo notebook project into one connected environment for continuous model and agent improvement across models, frameworks and clouds, with Agent Lens improving failure detection by 20% and fixing issues at half the cost, and Serverless RL training 1.4x faster at 40% lower cost. The loop from production behavior back to the next training run had been split across tools from different vendors with signal lost at every handoff; Forge attempts to bring that loop into a single environment. All figures are vendor-reported comparisons without stated control details; Canva, Capital One and MasterClass are named among the first companies building on Forge.
Perspective
This material is aimed at teams choosing AI infrastructure: AI labs, AI-native companies and global enterprises that need large-scale agent inference plus a post-training and evaluation loop can use these capabilities on CoreWeave Cloud through CoreWeave Kubernetes Service, SUNK, Mission Control, Sandboxes and Inference. The stated use cases include software engineering agents, reinforcement learning rollouts and model evaluation, as well as healthcare inference deployments such as Ennoble Care targeting clinical documentation, decision support and back-office automation.
All performance figures come from vendor or customer self-testing and are explicitly labeled early tests, so how far the 4.8x throughput, more than 3x sandbox startup, 1.7x Terminal-Bench and 1.4x training speedups hold across workloads and baseline configurations still needs more detail. Cognition's comparison reports only a relative multiple, with no task counts, absolute throughput or variance; Agent Lens's 20% failure-detection improvement and half-cost fixing also lack stated controls. In addition, this is a product and platform announcement narrative without a reproducible benchmark method or third-party evaluation, so readers making procurement or architecture decisions on this basis may want to wait for fuller benchmark disclosure.
