NeMo-DCR ships only the ~1% of weights that change, cutting 1T cross-cluster refit from 87.5 minutes to 150 seconds
Synopsis
NeMo-DCR introduces a bit-exact delta refit that transfers only the weights that change per step in canonical coordinates, using fixed affine mappings to project changes directly, residual conversion to cover the rest, mixed XOR/overwrite encoding, in-place application through the serving runtime's native loader, and a recoverable joint commit; in 3% and 5% stress cases across 30B–1T models, refits are 12–40 faster than a transport-only full-checkpoint reference, and a 1T relay-tree refit takes 150 s instead of 87.5 min.
Interpretation
Direct projection and residual conversion map changes from training shards straight into Hugging Face canonical coordinates, avoiding full-tensor assembly and conversion for affine changes, while residual conversion covers the rest, including changes from shared scale factors, against a distributed residual baseline. Prior systems either reimplement the serving runtime's placement rules or assemble and convert full tensors before extracting a delta; here affine changes are projected locally by shard owners and only residual tasks construct canonical tensors. On Qwen3-30B-A3B and Qwen3-235B-A22B at 3% and 5% stress rates, direct projection makes delta construction 1.08–1.16 faster than full conversion; affine mappings cover 96.3% of weight bytes in the Qwen3-30B-A3B checkpoint and 97.0% in Nemotron-3-Ultra-550B-A55B.
Mixed XOR/overwrite encoding keeps reconstruction bit-exact while improving compression: XOR masks carry changes on representation-preserving paths, and absolute overwrites carry the others. Arithmetic reconstruction can introduce floating-point rounding errors, and absolute overwrites resend unchanged bits within changed values; XOR masks flip only the bits that differ and therefore compress better. On a synthetic benchmark of 20 million BF16 standard-normal values with added Gaussian noise, XOR streams are 1.7–2.2 smaller than overwrite streams after zstd level-1 compression; on Qwen3, XOR encoding cuts payload bytes by 38–40% versus overwrites only.
Receivers leave placement to the serving runtime's native loader and apply updates in place without a receiver-side baseline copy; retries repair partial writes with overwrites, and a joint commit binds the policy version to its source baseline. In-place application avoids a separate receiver-side baseline copy, but an interrupted refit can leave a mixture of old and updated values; overwrite retries plus a joint commit recover from this, with a stated proposition and proof of bit-exactness. Comparisons with dense refits from the same candidate weights confirmed bitwise equality of every parameter and buffer element, including unchanged elements; in 50 GRPO steps that kill one vLLM instance every five steps, mean reward is 0.41 for dense NCCL and both NeMo-DCR transports.
Object storage or a relay tree streams payloads during delta construction without a cross-cluster collective, overlapping transport with construction. A cross-cluster collective couples receiver membership to the training cluster and introduces timeouts on failure; both transports here support deployments with and without direct links between clusters. Nsight Systems traces show the relay-tree transport lower bound accounts for 77–94% of refit latency versus 18–45% for the construction lower bound; 120B refits take 22.6–49.7 s and are 15.1–33.2 faster than the reference.
Perspective
The result targets agentic RL deployments that disaggregate rollout from training and must synchronize every policy update across a wide-area network, where the two sides choose parallel degrees independently and both support the Hugging Face canonical tensor layout. It supports both object-storage and direct-link topologies, covering deployments that share weights only through storage as well as those with direct links between clusters. The method relies on the serving runtime's native loader for placement and requires loader paths to be classified and validated in advance; unsupported loader paths are rejected ahead of time. The source baseline lives in host memory and totals about one checkpoint across training ranks, so applicability also depends on training-side host memory capacity.
Readers should still watch: the 3% and 5% change rates are stress cases above the measured 0.6–1.2%, and real change rates vary with learning rate and training stage, so gains will shift accordingly; the training experiment validates receiver kills on a 30B model over 50 GRPO steps, leaving larger models and longer runs to further observation; transport accounts for 77–94% of refit latency, so cross-cluster bandwidth and topology remain dominant factors; and this is a fast parse of the paper, so figures and appendix details are not fully represented, and reproducing specific configurations would require the original appendices.
