DMM has agents agree before they act: 1,598 of 1,600 MovingAI tasks solved, scaling past a million agents
Synopsis
The work identifies that in decentralized multi-agent path finding, even correctly learned per-agent action distributions can be recombined by independent sampling into incompatible joint actions, and proposes DMM, which replaces one-shot sampling with discrete iterative refinement of action intents across communication rounds before commitment, fine-tuned by MICPO, a critic-free multi-agent reinforcement-learning method; DMM generally achieves higher success rates and lower solution costs on POGEMA, solves 1,598 of 1,600 MovingAI tasks (the highest coverage among the evaluated methods), and scales to over one million simultaneously acting agents.
Interpretation
The paper identifies a structural limitation of learnable decentralized MAPF policies: even with communication, each agent still samples its final action independently after communication ends, so when several joint actions are valid in the same context, independent sampling recombines locally valid choices into collisions or deadlocks. Prior work (PRIMAL, MAPF-GPT, DHC, DCC, MAGAT/MAGAT+, HMAGAT, SCRIMP, LC-MAPF and others) enriches what each agent conditions on through reinforcement learning, imitation learning, and learned communication, but the final action is still sampled once and independently; this work attributes the failure to the sampling mechanism itself and gives a proposition on the decentralized factorization gap via conditional total correlation, showing the error is irreducible rather than a matter of imperfect learning. The paper provides a proposition with a proof in Appendix A and a controlled minimal-corridor experiment: two expert trajectories share one state where both resolutions remain possible, and DMM places 94.8% of its samples on the two valid joint actions, whereas LC-MAPF, MAGAT+, and HMAGAT sit near 50%, the independent-sampling outcome.
DMM reframes joint action selection as discrete iterative refinement of action intents across communication rounds: agents initialize intents from an uninformative prior, each round broadcast the current intent together with a learned message feature, sample a discrete vote, and update the intent, then each commits the action favored by its own final intent. Unlike diffusion or flow-matching refinement, DMM makes the intermediate intent itself part of the communication, so neighbors' stochastic choices can influence one another before commitment; unlike centralized refinement such as DiffLNS, which centrally refines the joint action tensor of all agents, DMM preserves decentralized execution with each agent relying only on local information and neighbor messages. The paper gives full equations for intent initialization, message broadcast, voting, and intent update, and instantiates two architectures: DMM-3M (3,241,784 trainable parameters, retaining the LC-MAPF Transformer encoder-decoder backbone) and DMM-0.8M (763,296 parameters); inference-time ablations in Appendices G and H show the evolving intent in communication carries more coordination-relevant information than the learned feature alone.
The paper introduces MICPO, a critic-free group-relative reinforcement-learning method that fine-tunes DMM's multi-round refinement on trajectory-level outcomes. Conventional actor-critic methods are poorly matched to decentralized MAPF: a centralized critic must generalize over a combinatorially large joint state, while a decentralized critic has only partial information; MICPO removes the value function, builds matched trajectory groups that share the sampled initial intent at each step, computes importance ratios per agent and per refinement round, and regularizes toward a frozen imitation-pretrained reference policy to limit drift. The paper defines the team return (a weighted sum of off-goal duration and blocked actions), group normalization, bounded replay, and the round-level clipped objective, and reports training configurations: both variants are imitation-pretrained for 1,000,000 iterations and MICPO fine-tuned for 500 outer iterations (96,000 optimizer updates), taking about 56.3 GPU-hours for DMM-3M and about 14.1 GPU-hours for DMM-0.8M on four H100 GPUs.
Across POGEMA, MovingAI, and a million-agent scalability experiment, DMM generally achieves higher success rates and lower solution costs, and attains the highest task coverage among the evaluated methods. On POGEMA, DMM-3M and LC-MAPF-3M share the same encoder-decoder architecture, communication bottleneck, and training data, differing only in how the final action is produced, which makes a controlled test of iterative intent refinement; on MovingAI, DMM operates reactively without constructing a search tree or revisiting executed decisions, yet outperforms search-based, hybrid, and learned baselines. At the maximum evaluated team sizes on POGEMA, DMM-MICPO-3M reaches success rates of 1.000 on Warehouse (192 agents) and 0.922 on Cities-Tiles (256 agents), versus 0.938 and 0.805 for LC-MAPF-3M, the highest-success baseline; on MovingAI, DMM-MICPO-3M solves 1,598 of 1,600 tasks and DMM-MICPO-0.8M solves 1,593, ahead of HMAGAT (1,576), LG-LaCAM (1,569), LaGAT (1,543), and MAPF-LNS2 (1,411); in the large-scale experiment DMM-MICPO-0.8M solves all 16 instances with up to 1,048,576 simultaneously acting agents.
Perspective
The result targets decentralized multi-agent path finding with local communication, partial observability, and the requirement that agents reach their goals without collisions, fitting settings such as warehouse fleets and city-scale autonomous transport and logistics. For practitioners, DMM offers a way to improve joint-action coordination while preserving decentralized execution: refinement rounds can be increased at inference without retraining, DMM-0.8M keeps the second-highest coverage at lower runtime, and DMM-MICPO-3M provides the best coverage and solution quality at a runtime comparable to MAPF-LNS2; under a 600-second budget, the median DMM-MICPO-0.8M and DMM-MICPO-3M runs would leave more than 99.6% and 98.6% of the budget for optional solution refinement. The paper also notes that DMM does not enforce joint-action feasibility as a hard constraint, so collision-free execution in the MovingAI and large-scale experiments still relies on an external mechanism such as CS-PIBT, and incorporating explicit feasibility constraints into refinement is listed as future work.
The corridor and POGEMA conclusions come from specific benchmarks and training configurations; on MovingAI the search-based and hybrid methods are limited to 600 wall-clock seconds while reactive policies are limited to 5,000 environment steps, so runtimes reflect solver configurations rather than an equalized compute budget, and the paper states that the decoding and RSE comparison describes complete inference configurations rather than a single-factor causal effect of argmax. In the large-scale experiment, shielding remains a substantial runtime cost, and at high density agents can still form one large conflict group that limits parallelism. In addition, the large-scale performance table (Table 1) appears with blank cells in this evidence bundle, so its specific numbers cannot be summarized here and only the qualitative trends described in the prose are available; readers needing exact step times, memory, and total runtimes should consult the original table.
