Cross-cohort evaluation of thyroid ultrasound deep learning across different outcome definitions: a duplicate-controlled benchmark
Synopsis
This preprint trained five backbones under duplicate-controlled image-level partitioning (387 thyroid ultrasound images and 1,332 fine-needle aspiration cytology blocks from 385 public cases, labelled by postoperative diagnosis) and applied the frozen models to TN3K (n = 1,228, dataset-provided benign/malignant labels) and to a DDTI endpoint derived from radiologist TI-RADS categories (n = 637), finding ultrasound ensemble AUROC 0.825 and ResNet-50 0.888 on TN3K versus 0.477 and 0.437 against the TI-RADS-derived endpoint, with a cytology benchmark AUROC of 0.986 (0.970–0.998) against 0.733 for the ultrasound ensemble, while noting that cohort, acquisition and endpoint changed together so the contrast cannot be attributed to label definition alone.
Interpretation
Exact MD5 hashing and perceptual hashing were used to group duplicates and near-duplicates before image-level partitioning, so that no detected duplicate group crossed the train, validation and test partitions (29 exact-duplicate rows and 33 same-perceptual-hash rows were detected originally, with three and four groups respectively crossing the original split). Relative to prior thyroid imaging AI work (Li et al., Peng et al., Duc et al. in Table 1, none reporting leakage safeguards), this study makes duplicate control an explicit part of benchmark design and reports the grouping statistics. Hash-based grouping statistics over 1,719 image metadata records, with cross-split group counts reported before and after regrouping; the authors state the procedure cannot establish patient-level independence because per-image patient identifiers were unavailable.
Frozen ultrasound models behaved differently across the two public cohorts: on TN3K the ensemble reached AUROC 0.825 (0.800–0.849) and ResNet-50 0.888 (0.870–0.904), whereas against the DDTI TI-RADS-derived endpoint AUROC was 0.477 (0.434–0.523) and 0.437 (0.391–0.482). The study explicitly separates evaluation against malignancy outcomes from agreement with an ultrasound risk-stratification system, rather than treating both as one external validation. External evaluation used no retraining and no threshold adjustment, with intervals from 2,000 bootstrap resamples; the authors note cohort, equipment, acquisition, preprocessing, disease spectrum and endpoint all changed together, so the contrast cannot be causally attributed to label semantics.
Within the development archive the cytology benchmark was more discriminative than the ultrasound benchmark: cytology ConvNeXt-Small AUROC 0.988 and the four-model ensemble 0.986 (0.970–0.998), versus 0.745 (0.612–0.865) for ultrasound EfficientNet-B3, with a descriptive non-paired AUROC difference of 0.247 (0.120–0.384). The two modalities are reported head-to-head within one archive, but the study states the contrast is non-paired and not patient-paired, so it is not evidence of clinical modality superiority. The cytology test partition held 200 images and the ultrasound partition 60 images, neither verified as independent patients; the difference used 5,000 independent resamples and is described as descriptive.
Two methodological add-ons give reusable operational signals: automatic texture/central ROI cropping raised ultrasound AUROC from 0.745 to 0.817 and improved Brier from 0.265 to 0.199, while temperature scaling (T = 2.414) left AUROC/AUPRC unchanged but reduced ECE from 0.0317 to 0.0138 and NLL from 0.2150 to 0.1262. Preprocessing and probability calibration are reported as separately testable steps, with the caveats that automatic cropping is not equivalent to expert nodule segmentation and that thresholds should be reselected after calibration. Both are exploratory analyses on the same small ultrasound test partition (n = 60) and cytology test partition (n = 200), positioned by the authors as hypothesis-generating.
Perspective
This benchmark applies to image-level, binary classification (benign versus papillary thyroid carcinoma) at 224 × 224 input resolution, with development labels from postoperative diagnosis and external evaluation using TN3K's distributed benign/malignant labels and a DDTI TI-RADS-derived grouping; its value is as a reproducible example of cross-cohort evaluation, duplicate control and probability calibration for imaging-AI researchers and reporting-guideline developers.
Readers should still watch that missing per-image patient identifiers leave patient-level independence unconfirmed, so internal estimates may remain optimistic; that public TN3K documentation does not establish pathology confirmation for every classification image; that the DDTI TI-RADS grouping is not a malignancy reference standard for all nodules; that cohort, acquisition and endpoint changed together so the label-semantics effect cannot be isolated; that automatic ROI cropping and Grad-CAM are exploratory and not specialist-reviewed; and that this reading was of incomplete scope, so the specific content of Figures 1–3 and the metric definitions in Appendix A could not be checked and may contain details relevant to interpretation.
