Public articles linked to the same research event.
arXiv The work proposes SCOUT, an offline multi-agent reinforcement learning framework that separately trains a flow-matching behavioral prior and a decomposed value function, then at test time transports behavioral samples toward high-value regions via Stein variational gradient descent, using the number of transport steps in place of a fixed regularization coefficient; under the individual-global-max (IGM) principle it proves a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation; empirically it achieves the best average performance across discrete and continuous offline MARL benchmarks and yields improvements in all offline-to-online configurations.
The work proposes SCOUT, an offline multi-agent reinforcement learning framework that separately trains a flow-matching behavioral prior and a decomposed value function, then at test time transports behavioral samples toward high-value regions via Stein variational gradient descent, using the number of transport steps in place of a fixed regularization coefficient; under the individual-global-max (IGM) principle it proves a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation; empirically it achieves the best average performance across discrete and continuous offline MARL benchmarks and yields improvements in all offline-to-online configurations.
The work proposes SCOUT, an offline multi-agent reinforcement learning framework that separately trains a flow-matching behavioral prior and a decomposed value function, then at test time transports behavioral samples toward high-value regions via Stein variational gradient descent, using the number of transport steps in place of a fixed regularization coefficient; under the individual-global-max (IGM) principle it proves a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation; empirically it achieves the best average performance across discrete and continuous offline MARL benchmarks and yields improvements in all offline-to-online configurations.
The work proposes SCOUT, an offline multi-agent reinforcement learning framework that separately trains a flow-matching behavioral prior and a decomposed value function, then at test time transports behavioral samples toward high-value regions via Stein variational gradient descent, using the number of transport steps in place of a fixed regularization coefficient; under the individual-global-max (IGM) principle it proves a single-term KL bound on the joint soft-value gap that vanishes as transport converges, with an irreducible additive residual proportional to the IGM violation; empirically it achieves the best average performance across discrete and continuous offline MARL benchmarks and yields improvements in all offline-to-online configurations.