Skip to main content
Back to timeline

An artificial intelligence-assisted diagnostic model for malignant head and neck tumors based on endoscopic images

Synopsis

This study retrospectively collected 23 434 electronic nasopharyngolaryngoscopic images from 3 255 subjects across five medical centers (15 465 laryngoscopic and 7 969 nasopharyngoscopic images), developed the WSC-T intelligent diagnostic model for head and neck tumors, reported internal and external test accuracies of 94.89% and 91.04% with AUCs of 0.98 and 0.97 for laryngeal-hypopharyngeal cancer and 96.27% and 92.31% with AUCs of 0.98 and 0.97 for nasopharyngeal cancer, compared it with a supervised learning model without contrastive learning, and built a cloud-based platform enabling image uploading, automated analysis, and diagnostic result output.

Interpretation

The study built an intelligent diagnostic model for head and neck tumors, named WSC-T, to identify malignant lesions of the nasopharynx, larynx, and hypopharynx. Relative to prior work, the model addresses both laryngeal-hypopharyngeal cancer and nasopharyngeal cancer tasks using electronic nasopharyngolaryngoscopic images, a commonly used clinical examination modality. Developed from 23 434 images (15 465 laryngoscopic and 7 969 nasopharyngoscopic) retrospectively collected from 3 255 subjects at five medical centers, and evaluated on internal and external test sets.

WSC-T reported high diagnostic accuracy and AUC on both internal and external test sets. The external test results provide direct evidence of cross-center generalization rather than internal validation at a single center alone. For laryngeal-hypopharyngeal cancer, internal and external accuracies were 94.89% and 91.04% with AUCs of 0.98 and 0.97; for nasopharyngeal cancer, accuracies were 96.27% and 92.31% with AUCs of 0.98 and 0.97.

The study compared WSC-T with a supervised learning model without contrastive learning. This comparison is intended to clarify the role of the contrastive learning component in the model design, rather than reporting only the absolute performance of a single model. The abstract states that this comparison was performed but does not give the specific numerical results of the comparison.

A cloud-based diagnostic platform was built on this model, enabling image uploading, automated analysis, and diagnostic result output. Packaging the model into an operable cloud workflow offers a convenient means for practical application of the intelligent diagnostic model. The abstract reports that the platform was successfully established with these functions, as a preliminary exploration of clinical feasibility.

Perspective

The study targets assisted diagnosis of malignant head and neck tumors (nasopharyngeal cancer, laryngeal-hypopharyngeal cancer) based on electronic nasopharyngolaryngoscopic images; the data were retrospectively collected from five medical centers, the model outputs image-level diagnostic results, and a cloud platform enables image uploading, automated analysis, and result output. Its conclusions apply to the internal and external test set evaluation settings described in the abstract and to the preliminary exploration of clinical feasibility.

The abstract does not provide the specific numerical results comparing the contrastive learning model with the supervised learning model, nor does it describe how internal and external test sets were split, the subject-level statistical basis, how image quality and acquisition device differences were handled, or the conditions and validation of the cloud platform in actual clinical workflows; these require the methods and results details of the original paper.

Sources