Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

SCOUT transports generative behavioral priors toward high-value regions via test-time action refinement, achieving the best average performance on discrete and continuous offline MARL benchmarks

The work proposes SCOUT, an offline multi-agent reinforcement learning framework that separately trains a flow-matching behavioral prior and a decomposed value function, then at test time transports behavioral samples toward high-value regions via Stein variational gradient descent, using the number of transport steps in place of a fixed regularization coefficient; under the individual-global-max (IGM) principle it proves a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation; empirically it achieves the best average performance across discrete and continuous offline MARL benchmarks and yields improvements in all offline-to-online configurations.