Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Finite-Depth Policy Sensitivity: Truncating Derivative Propagation Predicts Policy Adaptation to Interactive-Agent Behavior Change

The work develops a finite-depth framework that estimates policy sensitivity by approximating the policy Hessian and mixed derivative from reference-environment information, with an adjustable propagation depth determining where derivative propagation along the trajectory is truncated; the authors characterize the omitted derivative contributions and derive truncation-error bounds that are nonincreasing with depth and vanish at full-horizon propagation, and in a belief-driven pursuit-evasion game they find that derivative-estimation errors generally decrease with depth, that the method outperforms baselines in estimation accuracy and policy adaptation, and that sensitivity-based initialization improves zero-shot return over direct transfer and shows advantages for subsequent fine-tuning.