SlideDP speeds shared-host multi-GPU full-parameter fine-tuning by 1.46–2.64x and beats FSDP2's measured peak by 11.2% at a larger batch on four H100s
Synopsis
SlideDP introduces a synchronous data-parallel runtime for shared-host multi-GPU systems that maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks, achieving geometric-mean throughput ratios of 1.46–2.64 over SlideFormer, MegaTrain, and ZeRO-Offload in matched-batch sweeps, approaching GPU-resident FSDP2 at a smaller batch and exceeding its measured peak throughput by 11.2% at a larger batch on four H100s, while supporting 256K-token sequences for Qwen3-14B and fine-tuning Qwen2.5-72B on four RTX 4090 GPUs.
Interpretation
SlideDP extends host-resident layer streaming from a single-GPU workaround into a synchronous data-parallel runtime on a shared host: the host holds one authoritative copy of FP32 master weights and Adam states, each rank uses temporary BF16 layer copies in a bounded reusable GPU window, and a rank-0 optimizer worker updates that state once per iteration. Prior multi-GPU modes of SlideFormer and MegaTrain use replicated parameter delivery with CPU-side gradient reduction, so host demand grows with rank count, while partitioned offload systems such as ZeRO-Offload focus on reducing redundant state movement. SlideDP's difference is to decouple persistent state layout from communication routes while preserving synchronous DP semantics: the same parameter version across ranks, gradients aggregated across all ranks, and the next dispatch waiting for that update. The paper states three synchronous-DP invariants and per-layer parameter-version gating, and implements the runtime in PyTorch Distributed across RTX 4090, A800, and H100 platforms. For correctness, it compares SlideDP with GPU-resident FSDP2 over 200 training steps of Qwen3-14B on four H100s at per-GPU batch 16 and sequence length 1024, matching initial weights, C4 samples and their order, and optimizer settings; all ranks complete with finite loss and the maximum absolute differences in global-batch mean loss are reported.
The paper provides an analytical step-time model of shared-host scaling that uses shared-resource service demands and cross-layer, cross-rank execution dependencies to characterize host-bound and GPU-bound execution, explaining when reducing traffic shortens a step, when shrinking computation windows expose host work, and why the best route depends on both topology and workload. The model folds readiness, buffer reuse, and parameter-version dependencies into completion-time constraints and separates steady-state step time from the final update drain, giving a testable statement of why reducing logical byte counts alone is insufficient rather than an empirical tuning recipe. The model is presented as a capacity lower bound plus dependency completion-time equations and is examined through comparisons of communication routes, GPU counts, and chunk sizes; for example, on A800 removing NVLink bridges reverses the preferred delivery route in host-bound workloads, while route differences remain small in GPU-bound workloads.
AutoPolicy selects communication routes, chunk sizes, and activation layouts from measured pipeline effects, with Elastic Checkpointing supplying per-layer retention, offloading, and recomputation choices; on H100, Qwen3-14B throughput at batch 16 rises from 10,231 to 12,832 tokens/s (+25.4%), and policy selection gains 21.7% over the fixed policy at 64K. Unlike fixed checkpointing policies aimed only at memory reduction, this treats GPU memory allocation itself as a runtime policy decision: retaining activations removes offload/reload traffic while retaining intermediates reduces recomputation, and their benefit depends on which component is currently exposed. Table 3 reports per-configuration comparisons across platforms, for example a configuration on four H100s at batch 32 using 76.6 GiB versus 15.0 GiB for the fixed policy and gaining 12.4% throughput, and a configuration on four RTX 4090s at batch 16 retaining all 36 layer inputs and reducing host PSS by 14.5%. On selection overhead, one selection run on four H100s with Qwen3-14B at 32 sequences per GPU takes 1218 s, with a measured 0.970 s saving per step yielding an estimated break-even of 1,257 steps (2.72 hours of baseline training), or 1.4% of a 24-hour training budget.
In end-to-end evaluation, SlideDP achieves geometric-mean speedups of 1.83 over ZeRO-Offload and 1.48 over MegaTrain on four RTX 4090s, and 1.46 and 1.59 on four H100s; at batch 256 it reaches 23.2K tokens/s and 1,048,576 tokens per step without gradient accumulation, and it supports 256K-token sequences and Qwen2.5-72B fine-tuning on four RTX 4090 GPUs. These results translate host-memory capacity into larger feasible training workloads: on RTX 4090, Qwen3-32B reaches 4.24x ZeRO-Offload's throughput and Qwen2.5-72B reaches 240 TFLOPS while the other three offloading baselines are marked OOM; on H100, 4B–72B models sustain 1.8–2.0 PFLOPS. Evaluation covers three interconnect topologies (a PCIe-only workstation, an NVLink-bridged A800 server, and an NVSwitch H100 server); batch-scaling runs normally use 15 iterations discarding the first eight as warmup and report the arithmetic mean of two runs. For weak scaling, Qwen3-32B at 64 sequences per GPU maintains 95–98% parallel efficiency across two to eight A800s; for strong scaling, Qwen3-14B at global batch 256 achieves 1.94x scaling from four to eight GPUs versus 1.04x for SlideFormer.
Perspective
This work targets single-node, shared-host multi-GPU synchronous data-parallel full-parameter fine-tuning, with evaluation across three topologies: a PCIe-only workstation, an NVLink-bridged A800 server, and an NVSwitch H100 server, and model coverage from Qwen3-1.7B to Qwen2.5-72B plus two MoE models. It lets host-memory capacity translate into larger feasible batches and longer sequences, for example over 1M tokens per step for Qwen3-14B on four H100s, 256K-token sequence support, and Qwen2.5-72B fine-tuning on four RTX 4090s. For a reader, this means the feasible boundary of full-parameter fine-tuning is redrawn on nodes with limited GPU memory but ample host memory; AutoPolicy's selection overhead is 1218 s on four H100s in one run and is estimated to amortize after 1,257 steps, suiting settings with enough training steps.
The analytical model relies on measured effective service rates, which vary with the interaction of concurrent casting, DMA, and Adam, so how closely model predictions match a given hardware configuration still needs per-platform calibration. Selection overhead and amortization are reported for one run on four H100s with Qwen3-14B at 32 sequences per GPU, and the scale of that overhead for other models and topologies remains an open question. Evaluation covers dense models and two MoE models, and behavior across more architectures and longer training horizons is yet to be observed. In addition, this is a full-text reading, but some figures are presented as images, so specific plotting data and some numerical details require consulting the original figures.
