Public articles linked to the same research event.
arXiv The work introduces the Homomorphic Advantage Operator (HAO), which applies the zero-mean centering projection from advantage-based value estimation directly to temporal-difference targets to annihilate the uniform state-value baseline that drives Bellman drift; across a three-tier evaluation (tabular MDP, encrypted CartPole with real CKKS operations, and a 20-node logistics routing benchmark), HAO strictly bounds network pre-activations within the safe polynomial approximation domain with 0% boundary breaches across all random seeds, whereas regularization alone breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes, and HAO improves optimal policy accuracy by 18.0 percentage points in tabular domains.
The work introduces the Homomorphic Advantage Operator (HAO), which applies the zero-mean centering projection from advantage-based value estimation directly to temporal-difference targets to annihilate the uniform state-value baseline that drives Bellman drift; across a three-tier evaluation (tabular MDP, encrypted CartPole with real CKKS operations, and a 20-node logistics routing benchmark), HAO strictly bounds network pre-activations within the safe polynomial approximation domain with 0% boundary breaches across all random seeds, whereas regularization alone breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes, and HAO improves optimal policy accuracy by 18.0 percentage points in tabular domains.
The work introduces the Homomorphic Advantage Operator (HAO), which applies the zero-mean centering projection from advantage-based value estimation directly to temporal-difference targets to annihilate the uniform state-value baseline that drives Bellman drift; across a three-tier evaluation (tabular MDP, encrypted CartPole with real CKKS operations, and a 20-node logistics routing benchmark), HAO strictly bounds network pre-activations within the safe polynomial approximation domain with 0% boundary breaches across all random seeds, whereas regularization alone breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes, and HAO improves optimal policy accuracy by 18.0 percentage points in tabular domains.
The work introduces the Homomorphic Advantage Operator (HAO), which applies the zero-mean centering projection from advantage-based value estimation directly to temporal-difference targets to annihilate the uniform state-value baseline that drives Bellman drift; across a three-tier evaluation (tabular MDP, encrypted CartPole with real CKKS operations, and a 20-node logistics routing benchmark), HAO strictly bounds network pre-activations within the safe polynomial approximation domain with 0% boundary breaches across all random seeds, whereas regularization alone breached the bound on 3 of 5 seeds and the unstabilized baseline did so in 83.8% of episodes, and HAO improves optimal policy accuracy by 18.0 percentage points in tabular domains.