Securing Quantum Error Correction Against Misleading Advice from AI Agents
Synopsis
The work identifies an ambiguity in passive syndrome records that obstructs recovery selection and shows that in an odd-distance square toric code, opposite coherent X rotations produce identical passive syndrome-history distributions while a fixed phase correction helps at one sign and harms at the other; a terminal logical measurement on known encoded calibration states supplies the missing sign information, and a separate evaluator accepts an update only when calibration uncertainty and a justified drift bound certify improvement over the current recovery, thereby rejecting harmful proposals while retaining beneficial updates under honest advice in simulated advice attacks without assuming the adviser recommends correctly.
Figure 1: Trusted calibration separates advice from authority over recovery. (a) Opposite rotation signs give identical passive syndrome-history distributions in the odd-distance square toric instrument defined in Sec. II , although one fixed correction table can help at one sign and harm at the other. (b) Known encoded calibration states undergo 100 rounds of the same toric instrument and a terminal logical Y 03 Y_{03} measurement. Their sign-sensitive response supplies the missing observation; finite counts constrain the calibration angle. (c) The evaluator admits a proposed update only when sampling uncertainty and a justified drift bound certify recovery improvement at deployment; otherwise, the controller retains the incumbent or requests new calibration. The adviser can propose an action, including one induced by attacker-controlled text; the independent criterion determines whether that action may be applied. The displayed intervals are schematic. (d) The toric cycle uses ideal checks and X X recovery C s C_{s} ; the dashed V s ( u ) V_{s}(u) applies the specified logical operation without additional error. Circuit definitions are in Supplemental Material, Sec. V ; the encoded calibration protocol is in Sec. Y . Figure 2 gives quantitative evidence for the complete chain.
arXivInterpretation
Identifies an ambiguity in passive syndrome records: in an odd-distance square toric code with error-free preparation, syndrome measurements, and recovery operations, opposite coherent X rotations produce identical passive syndrome-history distributions, yet a fixed phase correction helps at one sign and harms at the other. Prior recovery selection typically relied on passive syndrome records; this work shows such records alone cannot distinguish the sign, clarifying the information gap required for recovery selection. Analytical argument on the toric code under error-free preparation, syndrome measurements, and recovery operations.
Proposes using a terminal logical measurement on known encoded calibration states to supply the missing sign information, and designs a separate evaluator that accepts an update only when calibration uncertainty and a justified drift bound certify improvement over the current recovery, without assuming the adviser recommends correctly. Shifts recovery updates from relying on adviser trustworthiness to relying on verifiable calibration evidence and a drift bound, yielding conditional guarantees. In simulated advice attacks, calibration-confidence checks rejected harmful proposals while retaining beneficial updates under honest advice.
Derives sufficient limits on calibration age that require improvement through deployment; in matched simulations, a validated channel-specific bound retains more beneficial updates than the general bound after accounting for evaluation time, while preventing the tested harmful activations under the stated drift assumption. Provides a quantitative relation between calibration information age and update acceptance, and compares retention between channel-specific and general bounds. Matched simulation results, contingent on the stated drift assumption.
In a surface-code experiment including stochastic circuit faults and noise changing during acquisition, deterministic controllers achieve at least as many beneficial updates with the same observations; violating the drift assumption permits harmful acceptance in the toric experiment. Extends the analysis from the idealized toric code to a surface-code setting with stochastic faults and noise drift, and notes the consequence when the drift assumption is violated. Surface-code experiment includes stochastic circuit faults and noise changing during acquisition; the toric experiment shows harmful acceptance when the drift assumption is violated.
Perspective
The results apply to an odd-distance square toric code under error-free preparation, syndrome measurements, and recovery operations, and to a surface-code experiment including stochastic circuit faults and noise changing during acquisition; for researchers and practitioners who want to safely adopt correction updates without fully trusting an AI adviser, it offers acceptance criteria conditioned on calibration measurements and a drift bound, along with sufficient limits on calibration age.
Violating the drift assumption permits harmful acceptance in the toric experiment, so the reasonableness of the drift bound in real deployment remains a point to watch; how the sufficient limits on calibration age and the channel-specific bound behave under broader noise models, and how to weigh the recovery improvements forgone through conservative acceptance, remain open directions.
