Skip to main content
Back to timeline
NVIDIA Technical BlogSource publication:

NVIDIA releases open source Topograph, turning cluster network topology into Kubernetes labels and Slurm configuration so schedulers place GPU workloads by physical proximity

Synopsis

NVIDIA introduces Topograph, an open source toolkit whose providers discover cluster topology from cloud APIs or on-premises fabric systems and normalize it into a canonical model, while engines translate that model into Kubernetes node labels, NFD resources, Slinky ConfigMaps, Slurm topology configuration, or topology JSON, letting schedulers on Kubernetes, Slurm, and Slinky make proximity-aware placement decisions and supporting topology-aware gang scheduling through KAI Scheduler and Kueue.

AI-generated editorial illustration: Topology-Aware Workload Scheduling with NVIDIA Topograph

Interpretation

Topograph unifies heterogeneous topology sources through two abstractions, providers and engines: a provider discovers topology from cloud APIs or on-premises systems and normalizes it into a canonical model, and an engine translates that model into the format each workload manager expects, including Kubernetes node labels, NFD resources, Slinky ConfigMaps, Slurm configuration, and instance-oriented topology JSON. Previously schedulers could only act on manually maintained topology snapshots; Topograph splits discovery, normalization, and format translation into replaceable layers so one topology view can serve Kubernetes, Slurm, and Slinky environments at once. The text describes this structure through a component list (API Server, Node Observer, Node Data Broker, Provider, Engine) and a provider-to-engine support matrix, which it states reflects upstream main as of September 16, 2026.

Topograph keeps the topology view current through both on-demand and event-triggered regeneration: the API exposes five endpoints (/v1/generate, /v1/topology, /v1/lookup, /healthz, /metrics), the Node Observer watches node or Pod changes and requests regeneration, and repeated identical requests reset a trailing timer and are processed once, with a typical aggregation delay of 15 seconds. It turns topology refresh from a manual snapshot into event-driven automatic regeneration and deduplicates requests during bursts of cluster events to reduce redundant work. The text gives endpoint semantics (submission returns HTTP 202, processing returns 202, completion returns 200) and the typical aggregation delay value, at the level of tool documentation.

On Kubernetes, Topograph expresses locality with a variable-depth fabric label family (fabric.topograph.run/tier-0 is the leaf switch closest to the node, with tiers increasing outward) and a two-level accelerator label (domain plus optional sub-domain); these labels serve directly as topologyKey values in Pod affinity and can be organized by KAI Scheduler into a hierarchy for gang scheduling. The default Kubernetes scheduler does not discover physical interconnect hierarchy, and Topograph fills that gap with node labels consumable by native affinity and topology-aware schedulers; in the KAI example the required annotation keeps the gang within a single tier-1 domain while the preferred annotation concentrates Pods in a tier-0 domain when feasible. The text provides label naming rules, verification commands, affinity weight examples (90 and 70), and complete YAML for the KAI Topology and Job annotations, making it a reproducible configuration example.

On the Slurm and Slinky side, Topograph generates tree, block, and the per-partition YAML configuration introduced in Slurm 25.05, with an optional reconfigure parameter that runs scontrol reconfigure after a file is written; the Slinky engine maps Kubernetes nodes to slurmd Pods and writes Slurm topology data to a ConfigMap, while the dra provider reads existing nvidia.com/gpu.clique labels for MNNVL block topology. It extends the same topology model to both native Slurm deployments and Slinky's Slurm-on-Kubernetes form, and covers the narrower MNNVL case. The text gives Debian/RPM build targets, the configuration file path, request and polling commands, tree/block/multi-topology JSON examples, and the note that the strigger script responds only to node up and down transitions and does not detect arbitrary switch rewiring.

Perspective

The toolkit targets cluster operators and platform teams running distributed GPU workloads on Kubernetes, Slurm, or Slinky, in both cloud-hosted and on-premises settings: on the cloud side a provider reads cloud APIs, while on-premises uses the InfiniBand provider with ibnetdiscover or NetQ for Spectrum-X and MNNVL domains. Prerequisites include Kubernetes 1.27 or later, Helm 3.10+ or 4.x, kubectl permissions, and a supported provider; topology-aware gang scheduling optionally uses KAI Scheduler or Kueue TAS. The NFD engine requires the alpha NodeFeatureGroupAPI feature gate, which is off by default, and its output is for downstream components consuming NodeFeatureGroup objects rather than a substitute for Kubernetes topologyKey labels. The text also provides kwok-nodes and Kind/KWOK helpers to test simulated node and switch hierarchies without production hardware.

The text explicitly notes that Topograph reflects reported rather than intended topology, that labels refresh only when generation runs, and that visibility of a fabric change depends on the provider and its triggering events; the support matrix is stated to reflect upstream main as of September 16, 2026, with requirements varying by Topograph version, environment, and provider configuration. The topology-aware workload scheduling introduced in Kubernetes 1.36 through KEP-5732 remains alpha, with upstream beta work ongoing, and the text advises consulting the enhancement tracker rather than depending on a specific future release. The strigger script for node-state refresh responds only to node up and down transitions and does not detect arbitrary switch rewiring or every inventory change. In addition, the text provides no quantitative measurements of throughput, cost, or tokens per watt, so those benefit directions are design intent rather than measured results in this article.

Sources