MoSDOT aligns noise with coordination modes via semi-discrete optimal transport, letting offline multi-agent distillation students retain the teacher's multimodal support
Synopsis
The work identifies a teacher-training failure mode in which standard flow-based teachers pair noise with replay targets independently, routing nearby noise samples toward conflicting coordination modes and producing samples between valid modes that the distillation loss propagates to students; it proposes MoSDOT, which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training, improving endpoint quality and routing consistency on controlled diagnostics and offline MARL benchmarks, and studies a shared-randomness variant to expose the residual gap intrinsic to strict-product execution.
Figure 1: Overview of source-to-mode routing. (Left) Standard Flow Policy: independent noise–target pairing can entangle nearby source samples with incompatible coordination modes, producing off-support fan mass. (Middle) MoSDOT: semi-discrete optimal transport partitions noise space into capacity-matched cells assigned to coordinated behavior modes. (Right) Route quality is measured by off-support avoidance ( fan-free endpoints ), consistency of source-to-mode routing ( clean route consistency ), endpoint accuracy to valid modes ( strict success ), and balanced mass across modes ( balanced coverage ). Full metric definitions are provided in Appendix E.6 .
arXivInterpretation
The paper identifies a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples are routed toward conflicting coordination modes and the teacher produces samples between valid modes. Prior work focused mainly on the distillation process itself, whereas this work moves attention upstream to how the teacher distribution is constructed before distillation begins, identifying teacher-side routing artifacts as the source. The claim is supported by controlled diagnostics: on the landmark diagnostic, flow matching reaches fan-free 0.857 and route consistency 0.856, while Drifting has the lowest route consistency at 0.653 and the largest Jacobian off-diagonal ratio, indicating its per-agent output is strongly shaped by other agents' noise.
The paper proposes MoSDOT, which summarizes multimodal replay into a finite mode support with prescribed capacities and partitions noise space via conditional semi-discrete optimal transport into Laguerre cells assigned to single modes before teacher training. Relative to independent pairing, this replaces arbitrary pairing with capacity-controlled source-region-to-mode coupling, using empirical replay frequencies as capacities and performing source assignment on the joint action as a whole rather than as a product of per-agent modes. On the landmark diagnostic MoSDOT reaches fan-free 0.983, route consistency 0.972, success 0.991, and balance 0.961, best or near-best on all four; its Jacobian off-diagonal ratio is comparable to flow matching and well below IMLE and Drifting.
The paper shows teacher-side improvement survives decentralized distillation: under the same student architecture, seed list, aggregation rule, and training budget, the MoSDOT-distilled strict-product student retains higher fan-free, success, and balance, whereas IMLE and Drifting students lose substantial endpoint quality through distillation. This links teacher geometry to student geometry, suggesting that improving the centralized source assignment can improve the support preserved by decentralized actors. On student-side diagnostics MoSDOT reaches fan-free 0.971, success 0.971, and balance 0.963, versus 0.784/0.558/0.621 for the IMLE student and 0.364/0.685/0.328 for the Drifting student.
The paper uses a shared-randomness variant to separate the strict-product gap from teacher-side assignment error: on the XOR target, even when the MoSDOT teacher concentrates mass on the two valid tuples, independent local sampling still mixes the two modes after distillation, while a shared mode-selection signal concentrates the student on the valid tuples. The variant lets agents correlate mode choice through a public signal without exchanging observations, actions, or messages, attributing the residual error to independent local sampling rather than teacher assignment failure. The appendix gives analytic results: a strict-product student matching marginals still assigns mass to unsupported tuples, while the shared-randomness policy, as a mixture of product distributions, exactly represents the anti-correlated XOR target.
Perspective
The result targets offline MARL researchers and practitioners learning coordinated policies from fixed datasets, in settings where replay admits a discernible finite mode support and where agent counts are modest with small discrete or low-dimensional continuous action spaces. The method performs dual optimization, Laguerre assignment, finite-cache rebalancing, and cache storage once before teacher training, while the deployed actor remains a one-step network, so inference cost is unchanged. The shared-randomness variant suits deployments that permit a public coordination channel, letting agents agree on mode selection without exchanging observations, actions, or messages.
The strict-product constraint is itself the central limitation: independent local sampling cannot represent arbitrary correlated joint support, and the shared-randomness variant exposes rather than removes this gap while requiring a public coordination channel whose admissibility and required shared dimension remain to be quantified. The empirical scope covers three offline MARL benchmarks with at most a modest number of agents and either small discrete or low-dimensional continuous action spaces; scaling to larger agent populations, richer action spaces, or online fine-tuning is not yet validated. In addition, the benchmark result tables in the provided text are empty, so specific SMACv1/v2 and MPE reward values cannot be checked here and only the trend descriptions in the main text can be relied on.
