Nereus adapts RL post-training parallelism on the fly: 27.7% lower average step latency on a real-data trace and up to 7.27x OpenRLHF throughput for 8B PPO
Synopsis
Nereus is a cost-aware runtime for reinforcement-learning post-training of large language models that represents each model-stage replica as an Elastic Model Unit and orchestrates Split, Merge, Extend, and Destroy primitives through a global transition graph, adapting the DP/TP/PP execution plan online under a payback rule; on a trace built from real data, online TP/PP adaptation cuts average step latency by 27.7% versus the initial fixed TP/PP layout, six transitions in a 1,000-step run scaling to 1,024 GPUs consume 0.079% of total run time, and end-to-end 8B PPO throughput improves by 2.14-7.27x over OpenRLHF and 1.10-1.47x over Verl.
Interpretation
A low-overhead, cost-aware adaptation policy: the controller monitors sequence length, available GPUs, peak memory, and achieved compute and communication efficiency; the planner selects the memory-feasible global plan with the lowest predicted steady-state step latency; and admission executes a transition only when the savings repay the transition cost before the triggering signal next crosses its threshold. Existing RL post-training frameworks fix allocation and parallelism after startup, elastic training systems typically manage a single model, and checkpoint-based resharding restarts the job and reconstructs state; Nereus applies the payback admission principle (in the spirit of Pollux and Sia scheduling) to one coupled RL job and separates an urgent path (profitability waived when the current plan is infeasible) from an opportunistic path. On Cluster #2, the cost model's training-stage latency term shows 4.96% mean absolute percentage error and 14.61% maximum error on 72 out-of-sample plans spanning 8-256 GPUs and four sequence-length buckets; decision overhead ranges from 0.17 ms at 32 GPUs to 338 ms at 1,024 GPUs, versus 3.4 ms to 511 s for a SCIP-based solver, with peak replanning memory below 100 MB.
A dependency-aligned state abstraction, the Elastic Model Unit (EMU): one EMU is one model-stage replica that encapsulates the tightly coupled TP/PP layout and state inside the unit while exposing loosely coupled DP replicas, so the DP degree is the number of active EMUs for that model-stage; online adaptation reduces to four primitives, Split, Merge, Extend, and Destroy. A coarse boundary (checkpoint or whole job) moves unrelated state, while a shard-level boundary exposes TP/PP resharding decisions to the planner; EMU picks the intermediate model-stage granularity so the planner reasons about whole units while primitives handle the underlying shards, preserving training semantics. On the same 1632-GPU 8B actor/critic workload, resource-scaling cost is 836.74 s for checkpoint/restart, 66.43 s at shard level, and 6.52 s for EMU; Extend takes 2.45-8.48 s from 48 through 3264 GPUs, 3.8-9.9x faster than Oobleck, 6.7-10.8x faster than Gemini, 9.3-16.2x faster than Tenplex, and 115.7-284.1x faster than UCP; for a 70B model on 64 GPUs, resharding cost drops 99.1% versus MCP and 96.7% versus Tenplex.
Safe concurrent transition orchestration: global plan changes are compiled into a directed acyclic graph of EMU primitives, with two deterministic rules, release-before-acquire and stage order, inserting cross-model-stage resource-dependency edges to avoid deadlocks on shared GPUs and letting ready primitives run concurrently across GPU groups and streams. Prior per-component migration does not handle transient GPU ownership conflicts; Nereus adds cross-stage edges only when transient GPU overlap blocks an acquisition and, when blocked, selects an executable Destroy that releases the needed GPUs, keeping the DAG acyclic and free of circular waits. Across 100 trials per transition type with randomized launch timing and GPU assignment, Nereus inserts 1-2 resource-dependency edges and succeeds in all trials, whereas DynaRL-style per-component migration succeeds in only 34-62%, with every failure a shared-GPU deadlock that a longer timeout cannot resolve; transition-DAG planning takes 1.22 ms at 1,024 GPUs versus 120.10 s for SCIP, with the transition completion-time gap to SCIP within 6.9% at every scale.
End-to-end validation: across three GPU clusters spanning vendors, interconnects, and scales, Nereus improves 8B PPO throughput while tracking training behavior, and the gains extend to ReMax, GRPO, and asynchronous execution. Relative to the fixed startup configurations of OpenRLHF and Verl, Nereus adapts within the same GPU budget as sequence length and stage bottlenecks change; relative to DynaRL-style admission, it additionally requires savings to repay the transition cost. 8B PPO throughput improves 2.14-7.27x over OpenRLHF (median 3.99) and 1.10-1.47x over Verl (median 1.21); selected plans stay within 5% of the empirical optimum in all 18 measured settings, exactly matching it in 61.1%; across three held-out traces Nereus averages 858.7 s/step versus 928.3 s for DynaRL-style admission; latency drops by up to 13.3% for ReMax and 15.8% for GRPO, and by up to 37.3% versus Laminar in the asynchronous setting; over the first 50 steps reward and PPO KL estimate track Verl with a 13.9% wall-clock reduction.
Perspective
This work targets engineering and systems teams running RL post-training (PPO, ReMax, GRPO, and similar) on GPU clusters, in settings where resource supply, sequence length, memory pressure, and stage bottlenecks change within a single run; adaptation happens at safe boundaries, namely RL-step completion in synchronous execution and weight synchronization in asynchronous execution. It lets a running job change its DP/TP/PP layout without reconstructing the full job state and replan from intact units after sudden node failures, complementing fault-tolerance systems. The authors note that the model-stage boundary and modular runtime provide a path to extend adaptation to context/expert parallelism, autoscaling external tool services, and additional RL frameworks.
The cost model stores one efficiency per operation class (for example GEMM, attention, or an all-reduce on one link type), and unseen classes start with conservative values refined after each step, so early decision quality on new models or hardware remains an open question; the plan space is limited to homogeneous TP/PP layouts within a model-stage and power-of-two parallelism degrees, and the authors list context/expert parallelism as future extension. Evaluation centers on Llama-3.1-8B with PPO, so behavior on larger models and more algorithm combinations still needs more validation; in addition, this evidence bundle is the full paper text with figures described in prose, so specific curve details cannot be checked point by point here.
