ENDOVISTA-ENT: Development and Preliminary Validation of an Integrated Artificial Intelligence System for Quality-Control Assistance and Lesion Recognition During Nasopharyngolaryngoscopy
Synopsis
In this retrospective two-center study of 2 365 patients (1 562 from the First Affiliated Hospital of Sun Yat-sen University and 803 from Ruijin Hospital, Shanghai Jiao Tong University School of Medicine), the authors developed ENDOVISTA-ENT, an integrated AI system with Model 1 for inside/outside-body image determination, Model 2 for recognition of 11 standard anatomical sites, and Model 3 for lesion localization and benign/malignant classification, reporting internal/external accuracies of 99.44%/99.84% and 96.09%/94.73%, malignant-lesion AUCs of 0.986/0.968, early nasopharyngeal, laryngeal, and hypopharyngeal cancer AUCs of 0.878-0.932, an increase in six physicians' overall interpretation accuracy on 200 pathologically confirmed cases from 78.50% to 88.20% with AI assistance (χ²=40.
Interpretation
The study built ENDOVISTA-ENT, a three-model integrated system covering examination quality control and lesion recognition: Model 1 determines whether images were acquired inside or outside the body, Model 2 identifies 11 standard anatomical sites, and Model 3 localizes lesions and classifies them as benign or malignant. Compared with single-task endoscopic AI tools, this work integrates quality control and diagnostic recognition into one system and reports internal testing and external validation performance on data from two independent centers. Based on 2 365 retrospective patients across two centers, with 1 562 for development and internal testing and 803 for external validation; Model 1 internal/external accuracy was 99.44% and 99.84%, Model 2 was 96.09% and 94.73%, and per-site AUCs ranged from 0.99 to 1.00.
Model 3 achieved high discrimination for malignant lesions and retained usable recognition performance for early-stage nasopharyngeal, laryngeal, and hypopharyngeal cancers. Beyond overall benign/malignant classification, the study separately analyzed early-stage upper aerodigestive tract tumors, a subgroup more prone to being missed clinically. Model 3 internal/external precision was 84.11% and 82.73%, recall was 80.60% and 78.31%, and malignant-lesion AUCs were 0.986 and 0.968; early nasopharyngeal, laryngeal, and hypopharyngeal cancer AUCs ranged from 0.878 to 0.932 across the internal test and external validation sets.
AI assistance improved lesion interpretation accuracy across physicians of different experience levels. Through a comparison in which six physicians of different experience levels performed independent and AI-assisted interpretation of 200 randomly selected pathologically confirmed cases, the study quantified the human-AI collaborative gain. Physicians' overall accuracy increased from 78.50% to 88.20%, χ²=40.37, P<0.001, based on 200 pathologically confirmed cases.
The system demonstrated real-time operating capability suitable for immediate feedback during examination. The study reported single-frame detection time, processing frame rate, and end-to-end latency, indicating the system is designed for real-time endoscopic scenarios. Model 3 single-frame detection time was (15.2±3.4) ms, the system processed (28.5±2.1) frames/s, and end-to-end latency was <40 ms.
Perspective
The system targets the nasopharyngolaryngoscopy setting, performing inside/outside-body image determination, recognition of 11 standard anatomical sites, and lesion localization with benign/malignant classification, and it can provide real-time feedback during examination. Its development and internal testing used 1 562 patients from the First Affiliated Hospital of Sun Yat-sen University, external validation used 803 patients from Ruijin Hospital, Shanghai Jiao Tong University School of Medicine, and the human-AI comparison used 200 randomly selected pathologically confirmed cases interpreted by six physicians of different experience levels. The study concludes that it provides a basis for further clinical validation, so the intended users are medical institutions and endoscopists performing nasopharyngolaryngoscopy with the relevant image acquisition conditions.
As a preliminary validation, the system's prospective performance in real clinical workflows, its robustness to image quality across different devices and operators, and the recognition difficulty suggested by early-lesion AUCs being somewhat lower than overall malignant-lesion AUCs remain questions for follow-up research; moreover, this text is at the abstract level, so specific model architectures, training details, and subgroup sample sizes are not elaborated, and readers needing methodological detail should consult the original article.
