Public articles linked to the same research event.
arXiv For non-stationary reinforcement-learning problems decomposable into phases, this work first shows that a single shared policy can theoretically achieve the performance of any multi-policy solution, then proposes a regime-based phase decomposition method that uses the duration of transient dynamics relative to the quasi-stationary period to identify which policy performs better, and validates four hypotheses across numerical experiments on different non-stationary RL problems: longer phase durations favor multi-policies, greater heterogeneity between phases increases the burden on a single policy, multi-policies need sufficient data for each phase, and environment-specific transition dynamics between phases affect which policy is preferable.
For non-stationary reinforcement-learning problems decomposable into phases, this work first shows that a single shared policy can theoretically achieve the performance of any multi-policy solution, then proposes a regime-based phase decomposition method that uses the duration of transient dynamics relative to the quasi-stationary period to identify which policy performs better, and validates four hypotheses across numerical experiments on different non-stationary RL problems: longer phase durations favor multi-policies, greater heterogeneity between phases increases the burden on a single policy, multi-policies need sufficient data for each phase, and environment-specific transition dynamics between phases affect which policy is preferable.
For non-stationary reinforcement-learning problems decomposable into phases, this work first shows that a single shared policy can theoretically achieve the performance of any multi-policy solution, then proposes a regime-based phase decomposition method that uses the duration of transient dynamics relative to the quasi-stationary period to identify which policy performs better, and validates four hypotheses across numerical experiments on different non-stationary RL problems: longer phase durations favor multi-policies, greater heterogeneity between phases increases the burden on a single policy, multi-policies need sufficient data for each phase, and environment-specific transition dynamics between phases affect which policy is preferable.
For non-stationary reinforcement-learning problems decomposable into phases, this work first shows that a single shared policy can theoretically achieve the performance of any multi-policy solution, then proposes a regime-based phase decomposition method that uses the duration of transient dynamics relative to the quasi-stationary period to identify which policy performs better, and validates four hypotheses across numerical experiments on different non-stationary RL problems: longer phase durations favor multi-policies, greater heterogeneity between phases increases the burden on a single policy, multi-policies need sufficient data for each phase, and environment-specific transition dynamics between phases affect which policy is preferable.