Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

Research Square

Cross-cohort evaluation of thyroid ultrasound deep learning across different outcome definitions: a duplicate-controlled benchmark

This preprint trained five backbones under duplicate-controlled image-level partitioning (387 thyroid ultrasound images and 1,332 fine-needle aspiration cytology blocks from 385 public cases, labelled by postoperative diagnosis) and applied the frozen models to TN3K (n = 1,228, dataset-provided benign/malignant labels) and to a DDTI endpoint derived from radiologist TI-RADS categories (n = 637), finding ultrasound ensemble AUROC 0.825 and ResNet-50 0.888 on TN3K versus 0.477 and 0.437 against the TI-RADS-derived endpoint, with a cytology benchmark AUROC of 0.986 (0.970–0.998) against 0.733 for the ultrasound ensemble, while noting that cohort, acquisition and endpoint changed together so the contrast cannot be attributed to label definition alone.