Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

A shared policy can theoretically match any multi-policy solution, but phase duration, heterogeneity, and per-phase data decide whether single or multiple policies win in practice

For non-stationary reinforcement-learning problems decomposable into phases, this work first shows that a single shared policy can theoretically achieve the performance of any multi-policy solution, then proposes a regime-based phase decomposition method that uses the duration of transient dynamics relative to the quasi-stationary period to identify which policy performs better, and validates four hypotheses across numerical experiments on different non-stationary RL problems: longer phase durations favor multi-policies, greater heterogeneity between phases increases the burden on a single policy, multi-policies need sufficient data for each phase, and environment-specific transition dynamics between phases affect which policy is preferable.