An XGBoost classifier cuts both missed and false ALMA bandpass anomaly flags
Related research and updatesSynopsis
This work frames ALMA bandpass calibration anomaly identification as supervised classification, using XGBoost over features built from Nadaraya-Watson kernel-regression scan statistics plus expert-informed features to learn non-linear decision boundaries; evaluated on 42,286 polarization-pair samples from Cycle 9 with execution-block-level nested cross-validation, it reduces both the false-negative and false-positive rates relative to the current ALMA pipeline subband heuristic on 19,567 applicable FDM polarization-pair samples, and shows consistent generalization on roughly an order-of-magnitude larger Cycle 11 dataset including the new Band 1.
Figure 1: Histograms of categorical features by bandpass solutions in the selected Cycle 9 training dataset.
arXivInterpretation
It presents an extensible supervised classification framework that turns bandpass amplitude anomaly detection from univariate thresholding into multi-feature non-linear classification, with the current implementation covering amplitude solutions only. Compared with reliance on human inspection or expert-crafted, manually tuned thresholds, the framework uses XGBoost to jointly optimize predictive performance across features and is explicitly designed to extend to phase features and other anomaly types. The framework, labeling definition, and feature table are given in full in the methods; training uses 84,572 Cycle 9 bandpass solutions (42,286 polarization pairs) with labels from QA2 flagging actions plus author review.
Core features come from Nadaraya-Watson kernel-regression scan statistics run in three ways: masked, unmasked, and fixed-width at 62.5 MHz, covering different anomaly situations. The masked version excludes known atmospheric line intervals to suppress false detections, the unmasked version handles anomalies overlapping atmospheric interference, and the fixed-width version matches the BLC FDM subband width to target platforming anomalies. Feature construction is itemized in Table 1; the authors also built an all-band atmospheric transmission model at 1 MHz resolution from the ATM code and used a peak-finding algorithm to label atmospheric interference regions.
On Cycle 9 data, adding features step by step raises the mean F1 across the five outer test folds from the scan-statistic-only baseline to the full feature set, with reduced fold-to-fold variation. Relative to the single scan-statistic baseline, adding interval width, unmasked and fixed-width features, and expert-informed features lowers both false-positive and false-negative counts in the confusion matrices. Evaluation uses nested 5x5 cross-validation with execution-block-level splitting, each outer fold serving once as test set so every labeled anomaly is evaluated exactly once; the training set is highly imbalanced, with anomalies a small fraction of the 42,286 polarization pairs, and the minority class is upweighted during training.
On the applicable FDM subset, the classifier reduces both the false-negative and false-positive rates relative to the current ALMA pipeline subband heuristic, and shows consistent generalization on Cycle 11 data including a new band. The authors attribute the improvement to two elements: XGBoost learns non-linear boundaries across multiple features, whereas the pipeline QA framework treats each feature independently with a univariate threshold; and the scan-statistic feature set is well matched to the target subband anomalies. The comparison uses 19,567 FDM polarization-pair samples and shows lower false-negative and false-positive rates; the Cycle 11 testbed contains 372,120 polarization pairs (744,240 solutions), where the classifier flagged 28 Band 1 pairs with 27 confirmed on review, and because only samples flagged by at least one method were reviewed, the reported rates are lower bounds.
Perspective
The result applies to bandpass calibration solutions in interferometric radio observations, with the current implementation covering amplitude solutions only; training and the main evaluation use public ALMA Cycle 9 PI science data, with labels defined from QA2 flagging actions plus author review. The intended setting is decision support for anomaly flagging in observatory pipelines: for subband platforming anomalies produced by the BLC in FDM mode, the classifier lowers both false-negative and false-positive rates relative to the current heuristic over 19,567 polarization-pair samples; the unseen-observation check uses 372,120 Cycle 11 polarization pairs, including Band 1 and 7-m array data correlated with the BLC, neither represented in training. Reusable outputs include the Cycle 9 training dataset released on Zenodo and an open-source repository with the model trained on the full feature set. The fixed-width feature is set at 62.5 MHz and stated to be adaptable to hardware, so the approach can transfer to next-generation wide-bandwidth correlator facilities with similar subband stitching.
The classifier still misclassifies in several situations: edge platforming anomalies are not fully evaluated because the scan statistics exclude the first and last channels, a single solution with two anomalous intervals conflicts with the single-interval assumption, and solutions with very few channels and high noise are hard to separate from random variation. The authors' review finds some false negatives come from propagated calibration artifacts caused by an anomalous reference antenna rather than missed subband anomalies, and whether those propagated artifacts are consequential can only be settled by assessing the calibrated data products. The Cycle 11 comparison does not treat QA2 flagging as ground truth and reviewed only samples flagged by at least one method, so the reported false-negative and false-positive rates are lower bounds; about three quarters of the solutions flagged by data reducers but classified negative relate to WVR LO leakage artifacts in Band 1, whose flagging policy has varied across cycles. In addition, this is a fast parse of the text in which equations, figures, and some table values could not be read in full, so details such as specific F1 values and confusion-matrix percentages should be checked against the original.
