BRAHMa detects galactic bars with YOLO11x oriented boxes, reaching 1.14 kpc mean absolute error on 3150 Hoyle galaxies, close to the 0.99 kpc disagreement between two human catalogues
Synopsis
The authors present BRAHMa, which converts Galaxy Zoo 3D volunteer bar masks into rotated-rectangle labels at a threshold calibrated against the Hoyle catalogue, fine-tunes a YOLO11x oriented-box detector on per-galaxy contrast-enhanced DESI Legacy images, and reproduces Hoyle bar lengths for 3150 galaxies with a mean absolute error of 1.14 kpc and Pearson 0.91 after a linear correction, close to the 0.99 kpc at which the two human catalogues agree, with no mask, threshold or fit at inference.
Figure 2: A GZ3D training galaxy. Left: the image. Middle: the volunteer bar mask, where brighter pixels were drawn over by more volunteers. Right: the mask superposed on the image, with the lowest-count pixels made transparent. The label box is derived from the middle panel alone.
arXivInterpretation
BRAHMa recasts bar detection from pixel segmentation into oriented-box regression: YOLO11x-OBB outputs the four box corners, from which centre, length, width and position angle follow, so no mask thresholding or post-hoc fitting is needed at inference. Earlier machine learning routes to bars were classifiers with saliency maps (Bhambra et al. 2022) or U-Net semantic segmentation (Walmsley and Spindler 2023; Cavanagh et al. 2024), leaving free parameters between mask and physical quantity; the authors state that, as far as they are aware, nobody had asked an oriented detector to find bars. On the held-out validation split precision is 0.972, recall 0.984, mAP50 0.993 and mAP50:95 0.695; on the test split precision is 0.968 and mAP50 0.990. The lower strict-overlap value is attributed to the intersection-over-union of two thin rotated rectangles falling quickly with small angle or length errors, and to scatter in the labels themselves.
The authors set an empirical floor for automated methods from the disagreement between two independent human catalogues: GZ3D-derived boxes versus Hoyle et al. (2011) lengths on 586 common galaxies give a mean absolute error of 0.99 kpc, a root-mean-square error of 1.40 kpc and Pearson 0.94 after a linear fit. This turns the 'ground truth' problem from an abstract caveat into a comparable number, so later model accuracy can be read against catalogue-to-catalogue disagreement rather than against a single catalogue. The figure is in-sample: one binarisation threshold and two fit parameters were tuned on the same 586 galaxies, and the authors state it is somewhat optimistic and slightly understates the true disagreement, with the size of that optimism not measured.
On the full 3150-galaxy Hoyle sample the detector's raw lengths are systematically larger by about a factor of 1.75, and after a two-parameter linear correction the agreement is a mean absolute error of 1.14 kpc, a root-mean-square error of 1.59 kpc, Pearson 0.91 and Spearman 0.92. The correction slope of 0.57 is close to the 0.545 slope between GZ3D boxes and Hoyle lengths, which the authors read as the detector having reproduced the GZ3D convention, so the linear correction is mostly a translation between two human conventions rather than a repair of a detector failure. The coefficients were first obtained by grid search on the matched subset and then rounded within the flat minimum of the error surface after inspecting the full sample; the authors note the correction was chosen with the same catalogue against which it is evaluated, but with only two free parameters fitted to more than three thousand galaxies they expect the optimism to be small.
Two data-preparation steps are described as new: every training image is enhanced with CLAHE whose parameters are chosen galaxy by galaxy by maximising the Fourier m=2 amplitude (bisymmetric strength), and the mask binarisation threshold is fixed by comparison with the independent Hoyle lengths. The enhancement acts only on the image presented to the detector, while the label boxes come from GZ3D masks and do not depend on the CLAHE parameters, so the selection cannot bias the labels; the threshold moves from a knob each analysis turns separately into a one-time choice inherited by every user of the trained weights. The threshold scan shows Pearson rising from 0.48 at 0.10 to a plateau of about 0.93 between 0.35 and 0.45 and falling above 0.5; the authors adopt 0.40, the centre of the plateau. Chosen enhancement parameters differ per galaxy, for example a clip limit of 2.3 for one and 5.0 for another.
Perspective
The tool is aimed at researchers who need bar lengths and position angles with one consistent definition across large samples, especially dynamical analyses of integral-field surveys and morphological statistics in imaging surveys. The authors note that for every barred MaNGA, SAMI or CALIFA galaxy, BRAHMa could supply the bar position angle and extent needed to set up the Tremaine-Weinberg integrals and to define the corotation radius, in a form that is the same for every galaxy and does not depend on who drew the mask, pushing the few-hundred-galaxy samples of Guo et al. (2019), Garma-Oehmichen et al. (2020) and Geron et al. (2023) towards the full barred population of the integral-field surveys. The detector can be run on every disc galaxy in the Legacy Surveys footprint to give bar lengths and orientations under a single consistent definition, and the same weights, or a light retraining, can be tried on Euclid and LSST imaging as it arrives, with higher-redshift rest-frame near-infrared JWST imaging a natural target. The authors also note that the per-galaxy contrast step is independent of the detector and may be useful on its own wherever a faint, roughly symmetric structure must be lifted out of a bright background before a machine learning model sees it. The detector runs through a public web interface where a user can upload a FITS cutout and receive the image annotated with the detected box.
Several open questions remain for a careful reader. How accuracy was assessed matters most: the mask threshold and the linear correction were both chosen using Hoyle et al. (2011) galaxies, and the 0.99 and 1.14 kpc mean absolute errors were computed on the same galaxies, so they are optimistic; the authors recommend selecting the threshold and fitting the correction on training folds and reporting the error on held-out folds. The training sample is drawn from MaNGA targets selected by stellar mass and redshift and keeps only galaxies in which at least five volunteers drew a bar, so weak bars, bars in dwarf galaxies and bars in highly inclined discs are under-represented, and the good agreement with Hoyle et al. (2011) should not be taken as evidence that the detector finds every bar in a flux-limited sample; a detection threshold study with a bar-free control sample is needed before measuring a bar fraction. Up to 586 of the 3150 Hoyle galaxies are also GZ3D galaxies and many were probably in the training split, and since the remaining galaxies were not evaluated separately, a contribution from training-set overlap to the quoted accuracy cannot be excluded. The candidate selection rule assumes the brightest pixel is the nucleus, so a saturated foreground star or a close companion's nucleus can pull the selection towards the wrong candidate. Lengths are projected with no inclination correction, while a dynamical application needs deprojected lengths requiring disc inclination and position angle from elsewhere. The detector has one class and returns one box, so double bars, bars with ansae and the boxy inner structure of near edge-on bars are probably collapsed into one rectangle or missed. The strict-overlap mAP of 0.70 suggests the box geometry agrees with the labels only to within the label scatter, some of which may be irreducible because a rectangle only approximates a bar's shape; the authors suggest more training data, test-time augmentation over rotations, or an ensemble that would also provide a per-galaxy uncertainty the present single model does not. Finally, the comparison with published methods is not like-for-like: BRAHMa's errors are measured after a linear correction fitted to Hoyle et al. (2011), whereas the Walmsley and Spindler (2023) figures appear uncalibrated and whether Bhambra et al. (2022) applied a correction was not established, and the Hoyle subsets used by each study are not guaranteed to coincide, so the correlation coefficients are the fairest basis for comparison.
