FAR-POLYP-SEG: A Prospective Single-Center Colonoscopy Dataset for Colorectal Polyp Segmentation with Patient-Level Metadata and Baseline Cross-Dataset Evaluation
Synopsis
This work prospectively collected 8,181 frames from 455 patients during routine colonoscopy at Farhikhtegan Hospital, Tehran, Iran, between February and December 2025 (432 polyp-positive frames with expert pixel-level segmentation masks and 7,749 normal-mucosa frames), linked patient-level metadata (age, sex, colonoscopy indication, BBPS score, and procedure duration) to every case, and trained and evaluated six segmentation architectures under one standardized protocol with patient-grouped five-fold cross-validation, finding that PraNet reached the highest internal Dice (0.755) and nnU-Net the highest internal IoU (0.665) and pixel accuracy, yet gated false-positive rates on normal mucosa ranged from 24.9% (YOLOv11m-seg) to 59.
Interpretation
It builds and publicly releases a prospective single-center colonoscopy polyp segmentation dataset containing 8,181 frames from 455 patients, of which 432 polyp-positive frames carry expert pixel-level segmentation masks and 7,749 are normal-mucosa frames, with patient-level clinical and procedural metadata (age, sex, colonoscopy indication, BBPS score, and procedure duration) linked to every case. Whereas many public colonoscopy datasets are limited in size, omit normal-mucosa frames, or lack linked patient-level clinical context, this dataset provides polyp-positive masks, a large set of normal-mucosa frames, and structured patient-level metadata as an imaging informatics benchmark resource. Prospectively acquired during routine colonoscopy between February and December 2025 without modifying the standard diagnostic workflow; frame counts, patient counts, mask composition, and metadata fields are explicitly stated.
It trains and evaluates six segmentation architectures (UNet, UNet++, UNet (MiT-B0), nnU-Net 2D, PraNet, and YOLOv11m-seg) under one standardized protocol using patient-grouped five-fold cross-validation, so that no patient contributes frames to both the training and the test set of any fold. By comparing multiple architectures under a unified protocol with patient-grouped cross-validation, it reduces evaluation bias from the same patient's frames appearing in both training and test sets and provides a reproducible baseline for the dataset. Six architectures, one standardized protocol, and patient-grouped five-fold cross-validation are explicitly described, with internal Dice and IoU metrics reported.
Internally, PraNet achieves the highest Dice (0.755) and nnU-Net the highest IoU (0.665) and pixel accuracy, but on the 7,749 normal-mucosa frames gated false-positive rates range from 24.9% (YOLOv11m-seg) to 59.5% (nnU-Net), almost completely inverting the internal Dice ranking, showing that the models most sensitive to real polyps are not the most specific. By adding a specificity evaluation on normal-mucosa frames within the same dataset, it reveals that internal Dice ranking alone can mask differences in false positives on normal mucosa, offering a sensitivity-versus-specificity perspective for validating segmentation algorithms. Internal Dice, IoU, pixel accuracy, and gated false-positive rates on normal mucosa are reported as concrete values, with 7,749 normal-mucosa frames.
All models are then evaluated on the public Kvasir-SEG dataset as an unseen external test set, with Dice ranging from 0.756 to 0.829 and the relative ordering largely preserved, although nnU-Net's internal advantage on accuracy and IoU does not carry over to external Dice. It provides a cross-dataset baseline evaluation showing that internal performance ordering is largely transferable to external data, while a given model's internal advantage on specific metrics does not necessarily extrapolate. External evaluation is performed on the public Kvasir-SEG dataset, reporting per-model Dice range of 0.756 to 0.829 and the preservation of ordering.
Perspective
The dataset is intended for validation of colonoscopy polyp segmentation algorithms and imaging informatics research, suited to internal benchmark comparisons and cross-dataset evaluation in a single-center routine colonoscopy setting; the released dataset, structured metadata, and evaluation code allow researchers to reproduce the baseline and extend it.
This is a single-center prospective dataset and external evaluation uses only Kvasir-SEG, so generalization across centers, populations, and devices remains an open question; whether the inversion between normal-mucosa false-positive rates and internal Dice ranking is stable under different data distributions, and what concrete gains patient-level metadata offer for algorithm development, both warrant further study.
