COMPASS replaces heuristic comet localization with YOLOv11 segmentation plus automated selection, raising bounding-box mAP from 0.664 to 0.713 on the DeepComet benchmark and cutting false positives from 130 to 6 across 15 out-of-sample images
Related research and updatesSynopsis
COMPASS is an open-source Python pipeline that replaces OpenComet's threshold- and geometry-based comet localization with YOLOv11-based instance segmentation and adds auditable automated comet selection with optional manual review; on the public DeepComet benchmark it raises bounding-box mAPbb@[0.5:0.95] from 0.664 to 0.713 and mask mAPmask@[0.5:0.95] from 0.584 to 0.593 on the hard subset, on 15 out-of-sample bat and painted turtle images it reduces false positives from 130 to 6 while increasing true positives from 128 to 146, its automated selection reaches 0.986 precision and 0.858 recall, and across 141 matched comets its olive moment agrees with OpenComet at a Pearson correlation of 0.945.
Figure 1: COMPASS pipeline visualisation.
arXivInterpretation
COMPASS replaces OpenComet's threshold- and shape-based comet localization with learned instance segmentation, achieving stronger localization on the public DeepComet benchmark: bounding-box mAPbb@[0.5:0.95] improves from 0.664 to 0.713, and mask mAPmask@[0.5:0.95] improves from 0.584 to 0.593 on the challenging hard-comet subset. Prior deep learning tools targeted segmentation accuracy with specific architectures, including DeepComet's Mask R-CNN with a keypoint head, GamaComet's Faster R-CNN, and a fully convolutional ensemble; COMPASS instead makes the default backend a YOLOv11 segmentation model fine-tuned on the DeepComet dataset while supporting alternative fine-tuned backends, turning localization into a swappable learned module. Evaluated with mAP on the expert-annotated public DeepComet dataset, reporting both bounding-box and mask metrics plus the hard subset; training and model details are in Supplementary Section S2 and full results in Supplementary Section S5 and Tables S1-S3.
On 15 fully held-out out-of-sample images (6 bat, Eptesicus fuscus, and 9 painted turtle, Chrysemys picta, spanning baseline, UV damage, and UV damage plus repair conditions), COMPASS reduces false positives from 130 to 6, false negatives from 27 to 14, and increases true positives from 128 to 146 relative to OpenComet, with detection precision rising from 0.636 to 0.964 and recall from 0.845 to 0.899. These images were excluded from model development, including training and model selection, and reference labels were established by independent review by members of the laboratories that collected the data, so the comparison is a transfer test across species, staining, and imaging conditions rather than a same-distribution benchmark reproduction. The out-of-sample evaluation covers 15 images with reference labels from independent review by the originating laboratories; painted turtle images used the default confidence threshold of 0.2, while bat images used an empirically determined threshold of 0.05 with the optional preprocessing step; per-image results are in Supplementary Section S6.
COMPASS provides an automated comet selection module that OpenComet does not offer: it assigns each comet a rule-based damage class from tail DNA percentage, then filters using treatment-specific criteria and quality flags, flagging overlapping comets, comets cut off at image borders, and low-signal or empty artifacts; on matched true-positive detections it reaches selection precision of 0.986 and recall of 0.858. Earlier deep learning comet tools improved localization and segmentation, but deciding whether a given comet enters downstream analysis still required expert intervention; COMPASS turns that user-facing decision into a rule-based, adjustable, auditable step, outputting green-box selected and red-box deselected visualizations plus a spreadsheet recording the factors behind each decision. Selection metrics are computed only on matched true-positive COMPASS detections to separate retain/reject decision quality from localization error; selection true positives/false positives/false negatives are 71/2/14, and the rules can be adjusted by users for different cell populations and laboratory processes.
On 5 additional painted turtle images with substantially denser comet populations (approximately 60 comets per image versus approximately 10 in the selection dataset), COMPASS agrees strongly with OpenComet on 141 matched comets, with olive moment showing a Pearson correlation of 0.945 and a concordance correlation coefficient of 0.857, while preserving the expected ordering of baseline, damage, and repair conditions. This analysis was used exclusively for measurement validation and these images were not included in the detection and selection evaluation; the result indicates that despite modest systematic differences in absolute values, both tools would support the same biological conclusions across experimental groups. 141 matched comets across 5 images, reporting Pearson correlation and concordance correlation coefficient and observing consistent treatment-dependent trends; full results are in Supplementary Section S7 and Table S5.
Perspective
The work targets researchers and laboratories using the single-cell gel electrophoresis comet assay who need standardized DNA damage measurements and reproducible selection decisions, and it applies to fluorescence or similar imaging conditions comparable to the DeepComet training distribution (canine peripheral blood mononuclear cells) or made comparable through the optional preprocessing step; for images that differ substantially, users can lower the detection confidence threshold, enable preprocessing, or swap the segmentation backend. The pipeline keeps a manual review interface that allows inspecting, accepting, rejecting, or adding detections, making it suitable for settings that need audit trails and cross-experiment consistency, such as compound genotoxicity assessment, disease mechanism studies, and cross-species aging research.
The out-of-sample detection and selection evaluation rests on 15 images (6 bat, 9 painted turtle), and measurement agreement rests on 141 matched comets across 5 high-density painted turtle images, so the scale is limited and behavior when extended to more species, staining methods, and imaging platforms remains to be observed. Bat images used an empirically determined threshold of 0.05 with preprocessing enabled while painted turtle images used the default 0.2, so how threshold and preprocessing choices affect results deserves further characterization. Modest systematic differences in absolute values were observed in the measurement comparison; these do not affect treatment-dependent trends but matter when comparing absolute values across tools. In addition, the text places some implementation details, training configurations, and per-image results in supplementary materials, so reading the main text alone does not allow full assessment of those aspects.
