QwenGyre pairs elastic GPU scheduling with trajectory-tree processing to lift Qwen 3.8 2.4T from 52.5% to 58.5% on NL2RepoBench in 48 steps
Synopsis
QwenGyre is an end-to-end framework for extreme-long-horizon (xLong) online agent RL that elastically reallocates GPUs between rollout and training without interrupting live harness executions, while its trajectory processor reconstructs branching executions into shared-prefix trajectory trees, scores partial progress after timeouts, and deduplicates sampled trajectories by role priority; scaled to Qwen 3.8 2.4T with 700K tokens per rollout, it raises NL2RepoBench pass rate from 52.5% to 58.5% in 48 steps and delivers up to 1.85x and 1.78x end-to-end speedups over Colocate and Async.
Interpretation
The paper identifies two structural difficulties for xLong online agent RL: severe execution-duration variance and prolonged rollout delays that leave GPUs idle, and complex non-linear branching that generates massive trajectory redundancy and cripples training efficiency. Prior work addresses scheduling or trajectory capture individually, whereas this work frames both as challenges that must be solved together in the xLong regime. Grounded in an xLong execution profile: a single execution can span hours, hundreds of model-environment interactions, and around 1M tokens across model calls; measured NL2RepoBench distributions for Qwen 3.6 122B show 1.93 hours mean per rollout execution, 2.96 hours per query, and 9.51% of queries taking at least four hours.
QwenGyre's elastic scheduler adjusts GPU allocation between rollout and training according to unfinished rollout work while keeping inference available to ongoing harness executions; newly available nodes can join an ongoing training batch, and streaming training distributes work according to each data-parallel group's progress. Unlike Async's fixed GPU pools or Colocate's pool-wide switching, this design keeps harness state alive across GPU role changes and lets training begin before the rollout tail completes. Under a 32-node budget with eight 4-node cells, E0 achieves a speedup over C2 with rollout fully hiding training and 20.85% slack remaining; average switching times are 8.52 s for rollout-to-training and 3.46 s for training-to-rollout, negligible against executions lasting hours.
The trajectory processor records exact input and output tokens plus behavior log-probabilities via token-in, token-out, organizes them into shared-prefix trajectory trees that preserve each output's original conditioning context, evaluates preserved workspace including assessable partial progress after timeout, and at admission selects at most a bounded number of trajectories per execution by role priority, counting shared targets once and averaging token losses within each execution. The design carries role, branch, environment-state, and evaluator provenance through rollout collection, admission, and training materialization, rather than addressing trajectory capture, credit assignment, or prefix computation in isolation. A case study reports one execution with 636 proxy requests, 1,347 unique message nodes, and ten candidate trajectories; independent expansion yields 1,716 node occurrences versus 1,347 shared nodes, with 369 extra occurrences repeating shared prefixes and system prompts. Ablations show sampling only the main trajectory yields lower training scores and larger gradient norms, while a cap of five matches sampling all trajectories and reduces forward-backward time by 25.2%.
Trained on Qwen 3.8 2.4T with 700K tokens per rollout for 48 steps, NL2RepoBench pass rate rises from about 52.5% to 58.5%; under equal GPU budgets and matched scheduling staleness, QwenGyre achieves up to 1.85x and 1.78x end-to-end speedups over Colocate and Async while matching baseline training scores. The result extends elastic scheduling plus trajectory processing to a flagship-scale model and xLong training workload, not only mid-sized configurations. Qwen 3.8 2.4T completes 48 NL2RepoBench training steps in 75.42 hours versus 134.55 hours for Async and 91.47 hours for Colocate; DeepSWE and TerminalBench with Qwen 3.6 122B over 24 steps report consistent speedups at two staleness targets.
Perspective
The framework targets xLong online RL settings where unmodified black-box harnesses run executions lasting hours, spanning hundreds of interactions and around 1M tokens, and it suits GPU deployments with multiple independently switchable cells and enough rollout capacity to serve unfinished work; at a fixed budget, eight 4-node cells versus four 8-node cells differ by only 2.2%, indicating that a modest number of cells retains most of the overlap benefit. For practitioners, this means idle GPUs can move from rollout to training while harness and workspace state persist, and assessable partial progress after timeout can still be used.
Open questions remain: efficiency under substantially smaller total GPU budgets is not evaluated; multi-step-burst streaming training cannot globally shuffle all minibatches before a burst starts, and the effect of sample ordering on learning quality is not isolated; the appendix's depth and time comparisons rely on a ready-rate approximation, and the counting identities concern group-weighted mean staleness, so worst-case and token-weighted guarantees need separate analysis; the case study supplies no designated task reward, so that execution's trainability can only be inferred from the admission rule.
