Skip to main content
Back to timeline
arXivSource publication:

Finite-Depth Policy Sensitivity: Truncating Derivative Propagation Predicts Policy Adaptation to Interactive-Agent Behavior Change

Related research and updates

Synopsis

The work develops a finite-depth framework that estimates policy sensitivity by approximating the policy Hessian and mixed derivative from reference-environment information, with an adjustable propagation depth determining where derivative propagation along the trajectory is truncated; the authors characterize the omitted derivative contributions and derive truncation-error bounds that are nonincreasing with depth and vanish at full-horizon propagation, and in a belief-driven pursuit-evasion game they find that derivative-estimation errors generally decrease with depth, that the method outperforms baselines in estimation accuracy and policy adaptation, and that sensitivity-based initialization improves zero-shot return over direct transfer and shows advantages for subsequent fine-tuning.

Source-provided article image: Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity
Fig. 1 ·

Fig. 1: Adaptation to changes in an interactive agent’s behavior by finite-depth policy sensitivity estimation.

arXiv

Interpretation

A finite-depth sensitivity estimation method is proposed that combines state sampling with local state-transition propagation to construct a Bellman-based derivative decomposition and a finite-depth approximation of the policy Hessian and mixed derivative, with adjustable propagation depth. Existing routes either rely on full-trajectory sampling, which demands substantial data, or on full-horizon model-based derivative propagation, which is computationally expensive; the method takes a middle ground by sampling only initial states and computing conditional expectations via the transition kernel and Bellman recursions. The paper provides the method derivation and the exact decomposition in Lemma 1, and evaluates estimates on a small-scale pursuit-evasion instance using spectral-norm relative error, showing the estimates are relatively insensitive to the number of initial-state samples.

Truncation-error guarantees are provided: the derivative contributions omitted by finite-depth propagation are characterized, and error bounds are derived for the approximated derivatives and the resulting policy sensitivity, with bounds nonincreasing in propagation depth and vanishing at full-horizon propagation. Prior work has mainly studied computation or approximation of policy Hessians, with less explicit treatment of structural error from truncating derivative propagation; this work quantifies the depth-accuracy tradeoff. Under the uniform derivative bounds of Assumption 2, Theorem 1 gives the error bounds and proves their nonincreasing and vanishing properties, with the proof unrolling the truncated Bellman recursion and counting omitted derivative terms.

In a belief-driven pursuit-evasion game, derivative-estimation errors generally decrease as propagation depth increases, the method outperforms baselines in both estimation accuracy and policy adaptation, and sensitivity-based initialization improves zero-shot return and benefits subsequent fine-tuning. Compared with a trajectory likelihood-ratio estimator, the method requires substantially less sampling data at the same sample count and consistently achieves lower error at several depth settings. A small-scale grid enables exact benchmarks for estimation accuracy, and a larger-scale grid evaluates zero-shot return and fine-tuning; across six transfers, sensitivity-based initialization improves zero-shot mean return over direct transfer, with the advantage persisting during fine-tuning in several cases and narrowing in others as both methods adapt.

Perspective

The framework targets adaptation settings where a locally optimal policy has been obtained in a single reference environment and the interactive agent's behavioral parameter changes by a small amount, and it applies to systems where the interaction structure is known or identifiable from data so that the reference transition kernel and its derivative with respect to the behavioral parameter are available. It lets practitioners trade model-based derivative propagation cost against approximation accuracy via an adjustable propagation depth, and use sensitivity as a first-order correction for initialization in the target environment, yielding a starting advantage in zero-shot return and subsequent fine-tuning.

The current analysis assumes access to the reference transition kernel and its derivative with respect to the behavioral parameter, and treats the boundary value estimate as parameter-independent, thereby truncating value-function derivative propagation; sampling error and value-function approximation error are not included in the theoretical bounds and are examined only empirically. Future work plans to relax the model-information requirements, incorporate sampling and value-function approximation errors into the analysis, and estimate boundary value derivatives to enable joint policy and critic adaptation. In addition, validation is concentrated in a single pursuit-evasion scenario, and generalization to other interaction structures and larger behavioral parameter changes remains an open question.

Sources