NVIDIA open-sources NVCRE: running real distributed workloads on Kubernetes turns pre-512-GPU-training cluster readiness into a provable property
Synopsis
NVIDIA released NVCRE (Cluster Readiness Engine), an open-source Kubernetes controller that runs real distributed workloads (NCCL communication, DCGM level-4 diagnostics, NeMo Nemotron 5 pretraining), measures results across topology-aware node groups, and reports exactly which nodes failed and why, verifying GPU cluster readiness before production workloads land instead of hand-writing NCCL manifests and bisecting racks manually.
Interpretation
NVCRE turns readiness from an assumption into a proven property: it runs real distributed workloads and reports which nodes failed each test, with a sample report showing communication/nccl-all-reduce failing after 42m 18s at full scale on 8 nodes, with gpu-node-07 flagged ThresholdViolation and gpu-node-12 HardwareFailureDetected. Previously platform teams encoded this progression in runbooks, spreadsheets, or shell scripts wrapped around NCCL tests, and often learned about degraded hardware from customer tickets; NVCRE packages the capability as a Kubernetes controller, removing the need to write NCCL manifests by hand or bisect racks manually. The text provides complete sample reports for both failure and pass cases (node names, failure reasons, runtime, bandwidth values) and states the built-in catalog covers three domains; it does not provide a controlled comparison against manual bisection or a statistical sample size.
NVCRE uses a three-resource CRD hierarchy—Certification, Workflow, Job—to attribute each failure to a specific node and category: a Certification names the nodes and categories, each category creates one Workflow, each Workflow creates its child Job, and results propagate upward. Compared with scattering test logic across scripts, this hierarchy groups results by category, allows inspection with kubectl, and fits GitOps workflows; the text gives the example of a run reporting that gpu-01 hit a hardware fault during NCCL and gpu-02 missed its bandwidth target. The text includes a Certification YAML example and kubectl commands and states the API is the product surface; this is design description and example rather than controlled evaluation.
For multi-node failures that cannot be attributed to any single node, testScale: diagnose runs topology-aware hierarchical group testing: the engine splits each failing group, reruns the halves until it reaches minGroupSize, flags groups that still fail as suspects, and maxConcurrent limits concurrent jobs so they do not saturate the fabric being measured. The text notes that when a 64-node all-reduce returns low bandwidth every node in the group is equally implicated and isolating the cause by hand can take days of engineering time; NVCRE outputs a small number of suspect nodes with the reason each failed instead of implicating the entire group. The text describes the algorithm steps and concurrency control parameter and the output form; it does not give success rates or timing data for this diagnostic mode on real clusters.
WorkloadRun automates the setup of multi-node GPU workloads: provide a container image, select a framework (exactly one of torch, mpi, or exec), and specify the number of nodes, and NVCRE generates the matching Kubeflow TrainingRuntime, injects the shared-memory volume, sets NCCL and platform environment variables, and enables NVLink scale-up networking where hardware supports it. The text notes Kubernetes has no built-in equivalent to a single Slurm srun command, so the same test requires GPU and RDMA resource requests, NCCL settings matched to the network fabric, a large enough shared-memory volume, and a way to ensure all pods start together; WorkloadRun fills these gaps and, as a plain CRD, can be used by external tools without adopting the rest of NVCRE. The text includes a WorkloadRun YAML example and framework field description, and explains that without a gang scheduler pods may deadlock and that spec.gangScheduler opts every pod into a gang-aware scheduler; this is mechanism description and example.
Perspective
The tool targets platform and operations teams running NVIDIA GPU clusters on Kubernetes, for readiness verification before production workloads land across bring-up, burn-in, preproduction, and production stages; it requires Kubernetes 1.29 or later, kubectl, Helm 3.x, and NVIDIA GPU Operator, with GB200/GB300 NVL72 catalog entries also needing the DRA Driver and the DCGM level-4 category needing the standalone DCGM service, and a gang-aware scheduler such as KAI Scheduler recommended on busy shared clusters. It records failed nodes and reasons but does not cordon, taint, or patch node conditions, so quarantine, drain, or external remediation requires pairing with the NVSentinel Certification Monitor.
The text gives no controlled comparison against manual bisection, no success-rate or timing statistics for the diagnostic mode, and no guidance on validating custom tests beyond the built-in catalog; thresholds ship with no defaults, and the illustrative busBandwidthGBps >= 900, goodputRatio >= 0.9, and avgTFLOPsPerGPU >= 800 values are labeled as examples for GB200 NVL72-class systems, so actual values must be chosen by the reader. The roadmap mentions support for new NVIDIA architectures, inference, and automated lifecycle validation but gives no timeline.
