Skip to main content
Back to timeline
RAS Techniques and InstrumentsSource publication:

BYOL features plus Protege active learning rank 100 MGCLS candidates, 99 showing diffuse radio characteristics and 55 confirmed as cluster-related emission

Synopsis

The work feeds self-supervised BYOL features extracted from source cutouts into the Astronomaly: Protege active-learning framework and evaluates the pipeline on high-resolution (about 7 arcsec), convolved (15 arcsec), and concatenated-feature MeerKAT Galaxy Cluster Legacy Survey (MGCLS) datasets, using tracers from a human-labelled catalogue as both guidance and benchmark; high-resolution features identify diffuse sources earlier than convolved ones, concatenated features perform best overall, and of the top 100 sources ranked by Protege 99 exhibit some form of diffuse radio emission with 55 confirmed as cluster-related, recovering 55 of 121 tracers from 62,587 sources with only 300 human labels.

Source-provided article image: A targeted machine learning approach for detecting diffuse radio emission with Astronomaly: Protege
Figure 2

Figure 2. Comparison of data cuts applied to MGCLS sources. Left: Gaussian component-based cut from Lochner & Rudnick (2024). Right: Beam-size cut proposed here. Plots show the number of sources (solid lines) and tracers (dashed lines) remaining in high-resolution (orange) and convolved (blue) versus cut threshold (Gaussian components or beam-size multiples). The vertical lines mark the selected thresholds. It is evident that the beam-size cut retains more tracers for similar subset sizes.

· Page 4

Interpretation

It develops and tests a targeted search pipeline that combines BYOL self-supervised features with Protege active learning to identify diffuse radio emission candidates in MGCLS. Earlier comparable approaches either rely on large labelled datasets or survey-matched synthetic simulations (such as Tuna and OpenCLIP-style methods), or use conventional anomaly detection that buries valuable sources in complex feature spaces; this work instead learns user preferences iteratively within a single dataset without predefining target morphology. PyBDSF extraction on MGCLS enhanced products yields 62,587 high-resolution and 40,879 convolved cutouts, reduced by a beam-size cut to 7,051 sources (121 tracers) and 4,319 sources (119 tracers); BYOL uses EfficientNet-B0 trained for 100 epochs with converging training and validation loss curves; PCA retains 95% of variance, reducing to 40 and 37 components.

The effect of resolution is quantified: high-resolution features recover tracers faster near the top of the ranking, convolved features recover more across a wider range, and concatenating both resolutions improves both early detection and overall recovery. Previous Protege applications did not systematically compare different resolutions of the same survey or test feature concatenation; this work provides cumulative-recall comparisons across four configurations. In the top 100, high-resolution features recover 54 of 121 tracers versus 55 for high-resolution concatenated features; in the top 500, high-resolution alone recovers 67 of 121 versus 75 when concatenated; convolved features recover 43 tracers in the top 100 versus 54 for convolved concatenated features, and 74 of 119 versus 75 in the top 500.

Visual inspection of the top 100 candidates shows the algorithm prioritises diffuse radio emission broadly rather than cluster-associated emission only: 99 show diffuse characteristics, 55 match known tracers, and the remaining 44 include radio AGN with jets or extended lobes, bent-tail and wide-angle-tail morphologies, face-on spirals with star-forming regions, and several unidentified extended structures. This reveals that radio morphology alone cannot separate cluster-associated diffuse emission from other extended structures, requiring X-ray or spectroscopic redshift information; the authors treat the 44 sources as candidates worthy of follow-up rather than failures. Based on visual inspection of the top 100 from the concatenated high-resolution coordinate subset, compared against the human-labelled catalogue of Kolokythas et al. (2025); cumulative recall curves are computed from tracers only and may therefore underestimate how well the algorithm highlights diffuse candidates.

The pipeline gives an actionable quantification of human labelling cost: after labelling 300 sources (20 iterations of 15), viewing 100 sources recovers 55 of 121 tracers, and viewing 250 recovers 61, a large reduction relative to the original 62,587 sources. The work raises the per-iteration labelling batch from 10 to 15 and adopts binary labelling (tracers scored 5, all others 0), trading flexibility for consistency, as a reference for exploratory searches at SKA scale. Results come from cumulative anomaly curves on the high-resolution and convolved datasets and the concatenated high-resolution coordinate configuration; the authors note diminishing returns in later iterations and a temporary drop in cumulative recall at the 180-label iteration relative to 120 labels for the high-resolution dataset.

Perspective

The result targets exploratory searches in radio surveys: when targets are rare, morphologically diverse, and cannot be defined in advance, a small number of human labels iteratively ranks candidates within a single dataset. It applies to interferometric surveys that already have high-resolution and convolved products, especially cluster samples like MGCLS; for new surveys with only one resolution or without known tracers, the beam-size threshold must be reselected and features retrained. The authors note that when both resolutions are available, concatenated features balance early detection and overall recovery; for rapidly locking onto the most prominent sources, high-resolution features alone suffice; for completeness, convolved features alone are preferable.

Tracers are used both to guide active learning and to evaluate performance, so cumulative recall is not a generalisation metric on an independent test set; whether the 44 non-tracer sources in the top 100 are new cluster-associated emission, diffuse AGN emission in galaxy groups, or background objects requires follow-up such as X-ray imaging or spectroscopic redshifts. The authors also note that binary labelling cannot express intermediate interest in ambiguous cases and that a gradient 0-5 scheme might provide a stronger learning signal; PyBDSF default parameters miss low-SNR sources and fragment very extended structures into multiple islands, affecting how candidates are represented; and cumulative recall for the high-resolution dataset temporarily drops at the 180-label iteration, suggesting that the balance of labelled examples in feature space affects ranking stability. In addition, this is a fast-parse version: the specific content of Figures 1 through 9 and Appendix Figures A1 and A2 is not included in the loaded text, so details tied to those figures can only be understood from the body text.

Sources