Online conformal calibration makes coverage steerable in RL-driven hardware-aware NAS, pruning 25-50% of evaluations at no measured accuracy cost
Related research and updatesSynopsis
The work replaces one-shot quantile estimation with online feedback control (Adaptive Conformal Inference with tuning-free, locally-adaptive, and group-conditional variants) to address non-exchangeability of candidates during reinforcement-learning search, tracking requested coverage to within about 10^-3 across three architecture families and both single-step and sequential search (three seeds) while pruning 25-50% of evaluations at no measured accuracy cost, whereas static calibration loses coverage control and a Gaussian-process baseline stays conservative.
Figure 1: Realized coverage vs requested target 1 − δ 1-\delta , per filter arm and family (points = seed mean, error bars = across-seed standard deviation; grey line = ideal calibration). The adaptive and tuning-free (agaci) filters lie on the diagonal with tight bars, calibrated and reproducible. GP-UCB is pinned near 0.99 regardless of the request. Calibrate-once split barely tracks the target (under-covering at tight requests, over-covering at loose ones), and its large error bars (absent from the means alone) are the reproducibility failure: biased and uncontrolled at once.
arXivInterpretation
One-shot quantile estimation in conformal-prediction filtering is replaced by online feedback control, making coverage monotonically and reproducibly steerable to a requested level for arbitrary sequences. The original approach assumes exchangeability between calibration and test candidates, which the RL loop violates because policy proposals improve as search proceeds and shift within every episode during layer-by-layer construction; online control removes reliance on that assumption. The abstract reports coverage tracking every requested level to within about 10^-3 across three neural-network architecture families, single-step and sequential search, and three seeds.
Online calibration prunes 25-50% of evaluations while maintaining coverage, with no measured accuracy cost. Static calibration loses control of its coverage and a Gaussian-process baseline stays conservative regardless of the request, so the online scheme achieves coverage control and evaluation-cost reduction together. The abstract states the pruning range and the no-measured-accuracy-cost result across three architecture families, two search modes, and three seeds.
Used as an acquisition function on one constrained testbed, the same optimistic bound beats random search. That gain is already carried by fixed optimism, and online calibration sharpens it, indicating the roles of calibration and acquisition can be separated. The abstract reports this comparison only on one constrained testbed, without specific values or effect sizes.
Perspective
The result targets hardware-aware NAS settings where the fraction of discarded candidates must be controlled during search, and it applies when the RL policy's proposal distribution changes as search proceeds and shifts within episodes during layer-by-layer construction. For researchers and practitioners who use conformal filtering to prune candidates and want coverage as a tunable parameter, it offers a way to dial the target coverage directly to the desired level, with reproducible code. Because the method addresses coverage control for arbitrary sequences, it is also informative for other filtering or screening pipelines that exhibit similar distribution drift.
At the abstract level, the text does not give per-family accuracy values, the distribution of pruning rates between single-step and sequential search, or the statistical uncertainty of the roughly 10^-3 coverage error; the acquisition-function comparison is also reported only on one constrained testbed, without systematic comparison to other acquisition functions. The abstract also does not describe the update frequency of online calibration, hyperparameter sensitivity, or behavior when early-search candidates are poor. These are open questions to watch in reproduction and extension.
