HAO applies zero-mean centering to temporal-difference targets, achieving 0% boundary breaches across all seeds in FHE reinforcement learning
Related research and updatesSynopsis
The work introduces the Homomorphic Advantage Operator (HAO), which applies the zero-mean centering projection from advantage-based value estimation directly to temporal-difference targets to annihilate the uniform state-value baseline that drives Bellman drift; across a three-tier evaluation (tabular MDP, encrypted CartPole with real CKKS operations, and a 20-node logistics routing benchmark), HAO strictly bounds network pre-activations within the safe polynomial approximation domain with 0% boundary breaches across all random seeds, whereas regularization alone breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes, and HAO improves optimal policy accuracy by 18.0 percentage points in tabular domains.
Figure 1: High-level overview of the HAO framework for privacy-preserving reinforcement learning. The system integrates three interdependent components: (1) CKKS-encrypted state transmission and polynomial forward evaluation, (2) the HAO centering projection that neutralizes Bellman drift, and (3) clipped weight updates with optional DP-SGD-style Gaussian noise.
arXivInterpretation
HAO applies the zero-mean centering projection from advantage estimation directly to temporal-difference (TD) targets, thereby annihilating the uniform state-value baseline that drives Bellman drift. Previously, applying FHE to reinforcement learning required replacing non-linear operations with polynomial approximations that diverge due to the recursive error phenomenon known as Bellman drift; HAO moves stabilization from the value-estimation level to the TD-target level. The abstract provides the mechanism: this linear projection maintains per-state action rankings while requiring zero additional non-linear multiplicative depth and avoiding expensive ciphertext bootstrapping.
HAO strictly bounds network pre-activations within the safe polynomial approximation domain, achieving 0% boundary breaches across all random seeds. Compared with regularization alone (L2 weight decay and gradient clipping) or an unstabilized baseline, HAO behaves differently on the boundary constraint. A three-tier experimental methodology: a tabular MDP, an encrypted CartPole environment using real CKKS cryptographic operations, and a 20-node logistics routing benchmark with dense continuous features; the comparison shows regularization breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes.
HAO agents improve optimal policy accuracy by 18.0 percentage points in tabular domains and remain stable when DP-SGD-style Gaussian noise is added to the clipped gradients. This extends the stabilization benefit from boundary constraints to policy quality and indicates compatibility with differential-privacy-style noise injection. The abstract reports an 18.0 percentage point accuracy improvement in tabular domains and states that stability is maintained when DP-SGD-style Gaussian noise is added.
Perspective
The work targets settings that require privacy-preserving reinforcement learning over confidential data in the cloud, applicable to intelligent systems where FHE replaces non-linear operations with polynomial approximations and where Bellman drift is a concern. Its three-tier evaluation covers a tabular MDP, an encrypted CartPole environment using real CKKS operations, and a 20-node logistics routing benchmark with dense continuous features, so the results apply directly to stabilization and policy quality in these settings. For readers seeking to deploy FHE reinforcement learning without additional non-linear multiplicative depth or ciphertext bootstrapping, HAO offers a path that applies zero-mean centering to TD targets and can be combined with DP-SGD-style Gaussian noise on clipped gradients.
Readers should still watch whether HAO maintains 0% boundary breaches beyond the tabular MDP, encrypted CartPole, and 20-node logistics routing benchmark at larger scales or across more task families; how well the zero-mean centering projection preserves per-state action rankings under different network architectures and polynomial approximation orders; and how the strength range of DP-SGD-style Gaussian noise relates to stability. The reading scope here is the abstract, which does not include the body figures and full experimental details, so these questions remain open at the abstract level.
