Skip to main content
Back to timeline
arXivSource publication:

WISE-ATTA shifts active test-time adaptation from which sample to label to when to label, reaching lower error on ImageNet-C/R/K/A with fewer labels

Synopsis

The work introduces budgeted active test-time adaptation (ATTA), in which labels are available for only a fraction of batches in a long test stream, and proposes WISE-ATTA: lightweight online signals (the fraction of low-entropy predictions) plus a sliding-history quantile threshold and a budget-debt controller decide when to request supervision, while prediction drift relative to an EMA anchor selects a single sample to label; on ImageNet-C under CTTA/FTTA and on ImageNet-R/K/A it matches or improves on recent ATTA methods at one label per batch and reports up to 50% fewer labels.

AI-generated editorial illustration: WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation

Interpretation

The paper reformulates active test-time adaptation as a budgeted problem: labels cover a fixed annotation ratio of test batches, exactly one label is queried from each selected batch, and the total budget is enforced by a constraint. Prior ATTA methods implicitly assume supervision can be requested for every incoming batch and focus on reducing labeled samples within a batch (SimATTA selects by entropy and clustering, EATTA by prediction sensitivity to feature perturbations); this work decouples supervision from batch structure, shifting the central question from what to label within a batch to when to label over time. The setting is formalized with a binary action and a total-budget constraint and evaluated on ImageNet-C/R/K/A; the text argues it better reflects practice where human, computational, or latency constraints make supervision only intermittently available.

WISE-ATTA's batch selection uses the fraction of low-entropy predictions as a proxy for batch-level adaptation reliability and requests supervision when the current score exceeds an adaptive quantile threshold over a sliding history, with the quantile level governed by a budget-debt controller that does not require knowing the stream length. Compared with uniform fixed-cadence labeling, the policy concentrates supervision on batches that are informative relative to recent observations; the authors report that budget-paced selection generally yields the largest error reductions on ImageNet-C under both CTTA and FTTA, with average improvements on ImageNet-R/K/A for both backbones. Evidence comes from controlled experiments that hold the sample selector fixed and vary only batch selection against Uniform and Random, plus full per-corruption results in the appendix; Appendix D shows average error varies only modestly across tested history-window, warmup, and slack values.

Sample selection uses prediction drift relative to an EMA anchor model: the single sample with the largest drift is queried from a selected batch, and the paper interprets drift as the trajectory sensitivity of a sample's prediction along the model's recent adaptation direction. The paper argues entropy is a static single-instant criterion that cannot distinguish ongoing adaptation from a stable end state, whereas drift captures whether predictions are still changing in a structured way; the query-behavior analysis reports Max-Entropy concentrating on near-zero-drift extreme high-entropy samples and EATTA largely in a low-drift band, while WISE-ATTA selects a moderate-entropy, higher-drift regime. At the same one-label-per-batch budget, WISE-ATTA achieves the lowest average error on ImageNet-C (CTTA 53.3%, FTTA 51.6%) and outperforms the three-label HILTTA; on ImageNet-R/K/A it reaches 70.9% average error for RN50-BN and 56.6% for ViT-B-16.

The paper reports the temporal shape of budget allocation and deployment-oriented analyses: label density is higher early in the stream and the budget is not exhausted early; advantages persist when updates are skipped on non-selected batches, across online batch sizes, and under label delay. These analyses ground the effect of when to supervise in observable allocation behavior and mark a boundary under delay: all methods degrade systematically as delay grows, with ViT-B-16/ImageNet-K error rising from 58.0 at delay 0 to 93.0 at delay 200. Evidence comes from counting labeled batches in ten equal time bins, the lightweight variant comparison in Appendix F (WISE-ATTA average error moving only from 55.7 to 55.8), the batch-size sweep in Appendix G, and the delay comparison table in Appendix H.

Perspective

The result targets image classification deployments with continuously drifting distributions, limited annotation budgets, and only intermittent access to supervision, and suits practitioners who want to cut labeling without adding a replay buffer or teacher model; the authors note the method remains competitive when unsupervised updates are skipped on non-selected batches, making it usable in compute-constrained lightweight deployment, and they report behavior across different online batch sizes.

The paper assumes each batch is dominated by a single distribution shift, and the authors identify heterogeneous batches with co-occurring shifts as a natural extension; the delay analysis shows all ATTA methods degrade systematically as delay grows, leaving delay-aware update weighting or latency-aware batch selection open; in addition, several tables in the main text and appendices are not fully rendered in the loaded text, so item-level comparisons should be checked against the original.

Sources