X-ray baggage screening drops manual labels: synthetic overlapping scans train a localizer, lifting mAP 2% to 23% in clutter
Lead
X-ray security scans lack human annotations, so LAO-X synthesizes overlapping baggage from unlabeled scans to train a SAM2 localizer, raising mAP by 2% to 23% over SAM2 and X-ray-specific baselines in cluttered scenes across six benchmarks.
Story
Object localization in X-ray security scans can now be trained entirely without human annotations, with the model finding objects of any category directly on real scans without category names or text prompts. Previous X-ray detectors either trained fully supervised on predefined categories or relied on category names or visual exemplars at inference, while X-ray annotation is costly and covers few categories, leaving models brittle to unknown objects and scanner shifts. Across six X-ray benchmarks, LAO-X improves mAP by 2% to 23% over SAM2 and X-ray-specific baselines in cluttered scenes, reaching AP50, AP75, and mAP of 63.39%, 55.73%, and 53.33% on DET-COMPASS.
The pipeline first mines salient objects from unlabeled scans, then recomposes them into annotated synthetic baggage images following X-ray transmission physics, and fine-tunes a SAM2 that emits masks at a requested granularity. Direct zero-shot transfer of SAM-family models pretrained on natural images is unreliable, because X-ray object boundaries come from density and projective superposition rather than visible surface texture. Training uses 1621 scans from COMPASS-XP to generate 18000 synthetic scenes, with an easy-to-hard three-stage curriculum that raises object count and overlap, while all six evaluation datasets are test-only.
What to watch
Next steps include validation and calibration on the target scanner, and using LAO-X as an assistive localization front end feeding downstream recognition and human inspection; for single-energy scanners the synthesis reduces to single-channel density-map composition.
Low-attenuation objects remain hard, with AP50 of 26.69% on the DET-COMPASS invisible subset versus 63.39% on the full set; under extreme overlap AR10 is 14.60, below SAM3 at 31.35. The synthesis does not explicitly model scattering, beam hardening, detector noise, or scanner-specific pseudo-color differences, so cross-scanner robustness still needs confirmation on target devices.
