SPL raises task completion in all 60 staggered-participation MARL comparisons, averaging a 14.1% gain
Synopsis
The authors formalize staggered participation (SP), a cooperative MARL setting where agents contribute to one objective over temporally offset participation windows and earlier participants leave task-relevant information for later ones, and propose SPL, a training-time augmentation combining prospective acquisition supervision for earlier agents with outcome-supervised receiver learning for later agents; across 60 MPE/RWARE backbone comparisons SPL shows higher observed mean task completion in every case, averaging a 14.1% difference, with gains extending to eight-agent teams and a physics-based UAV–UGV setting in Isaac Lab.
Interpretation
The paper formulates staggered participation as a learning dependency in cooperative MARL: an early action affects the return through the information it provides and through the later policy that uses it, so learning must address both what information to acquire and how later agents should use it. Prior work studies asynchronous interaction, dynamic or open teams, and communication-based information exchange separately; here information acquisition and downstream utilization are placed within one cooperative learning problem, with participation windows that may partially overlap or be fully separated. The setting is formalized as a Dec-POMDP with half-open participation intervals, where inactive agents take a null action and make no policy decision; it is instantiated in MPE, RWARE, and Isaac Lab.
SPL is a training-time augmentation that gives earlier participants prospective acquisition supervision and later collaborators outcome-supervised receiver learning, leaving the environment reward and backbone critic objective unchanged. The acquisition target is built from a fixed structured task model, while the receiver target is built from evidence-conditioned success probabilities learned from attributable outcomes combined with fixed task-level utilities, so no predefined correct-action labels are needed and no execution-time communication is added. Ablations show Acquisition-only and Receiver-only reduce mean task completion by 10.7% and 7.7% on average; SKD-MAPPO trails full SPL by 11.1% on average; Evidence-redacted, Outcome-agnostic, and Permuted reduce it by 10.7%, 18.0%, and 48.4% on average.
Across the 60 primary Vanilla–SPL comparisons, SPL shows higher observed mean task completion in every configuration, with an equal-weight descriptive mean difference of 14.09%; backbone-level differences are 13.93%, 14.02%, and 14.33% for IPPO, MAPPO, and HAPPO. The gains span five participation patterns, Base and Advanced environment variants, and three PPO-based backbones, indicating the improvement is not specific to a single actor–critic backbone. Four-agent experiments use five independent training seeds per method and configuration with 200 final evaluation episodes per seed; 57 configurations have paired inference, 56 pointwise 95% confidence intervals have positive lower bounds, and 29 remain significant after Holm correction at familywise level 0.05.
Gains extend to eight-agent teams and to a physics-based UAV–UGV setting in Isaac Lab, where earlier UAVs acquire evidence and later UGVs execute downstream tasks under physical motion and obstacle constraints. This provides evidence across algorithmic, temporal, and embodied settings, and SPL requires no acquisition model, outcome replay, outcome model, or utility estimate at execution time. SPL-MAPPO is higher in all six matched eight-agent comparisons with an average observed gain of 17.1%; Isaac Lab gains range from 20.6% to 30.1% across four team-size–participation settings; four-agent Isaac settings show roughly 14–16% more observed training time.
Perspective
The results apply to cooperative tasks where structured task knowledge is available during training: validity priors, the fixed valid-task constraint, the probe observation model, and task utilities come from the training environment or task specification, and all SPL-specific models and targets are training-only, with agents acting solely through their decentralized policies at execution. Applicable settings include multi-robot cooperation with partially overlapping or fully separated participation windows where earlier participants leave persistent task-relevant evidence, such as the Isaac Lab UAV–UGV task where earlier UAVs sense and later UGVs execute. Reported results characterize within-configuration performance, since all configurations are trained and evaluated independently, rather than transfer across participation patterns or task structures. The theoretical analysis characterizes local properties of the SPL auxiliary objectives, supporting construction and interpretation of the training signals without forming an end-to-end guarantee or establishing global convergence of the cooperative return.
The prior-misspecification experiment perturbs only the validity priors used by the acquisition model while the fixed valid-task constraint, probe accuracy, and task utilities remain correctly specified, so robustness to arbitrary errors in the full structured task model is not established. The statistical analysis is retrospective, uses only five paired seeds per configuration, assumes seed-level differences are approximately normal, and reuses a fixed 200-episode reset-seed schedule across training seeds and methods, so the inference is conditional on that evaluation design. In the four-agent Isaac setting SPL-MAPPO shows lower map coverage and task discovery than MAPPO while achieving higher task completion, and the reported diagnostics do not causally separate more selective acquisition from more effective use of available evidence. Several table cells in the loaded text are empty, so specific values require the appendix tables.
