Multimodal Large Language Models for Bladder Tumor Detection in Cystoscopy: A Retrospective Benchmarking Study
Synopsis
This retrospective study analyzed 1,754 labeled public cystoscopy images to test Direct, Book-based, and Optimized prompts across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano for benign-versus-malignant classification, finding that GPT-5 and GPT-5-Mini with the optimized prompt reached accuracies of 86.7% and 89.2%, and that GPT-5 with the optimized prompt achieved 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity in high-confidence triage at 62.0% image coverage, while prompt engineering improved calibration and triage without statistically significant performance gains.
Interpretation
The study applied multimodal large language models directly to benign-versus-malignant classification of cystoscopy images, reporting accuracies of 86.7% and 89.2%, specificities of 94.1% and 88.4%, and sensitivities of 82.5% and 89.2% for GPT-5 and GPT-5-Mini with the optimized prompt. Compared with existing AI approaches that often depend on data-intensive models difficult to deploy in routine practice, this work evaluates smaller and more efficient architectures, making deployability part of the assessment. A retrospective analysis of 1,754 labeled public cystoscopy images reporting accuracy, sensitivity, specificity, and F1 Score.
The study introduced high-confidence triage with an abstention option based on a loss function, where GPT-5 with the optimized prompt reached 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity at 62.0% image coverage. Beyond conventional classification, it adds an abstention-capable triage setting and reports coverage alongside accuracy, giving a quantitative description of which images can be read automatically. Triage evaluation on the same retrospective dataset, explicitly stating the 62.0% coverage condition.
The study compared Direct, Book-based, and Optimized prompts, finding that prompt engineering improved model performance, although these gains were not statistically significant, and enhanced confidence calibration and triage performance. It treats prompt design as an actionable variable and characterizes confidence reliability with Brier Score and Expected Calibration Error rather than classification accuracy alone. Comparison of three prompt types across four models, with calibration measured by Brier Score and Expected Calibration Error.
Perspective
The work targets benign-versus-malignant classification of cystoscopy images, using 1,754 labeled public images in a retrospective analysis; its conclusions apply to benchmarking prompt-and-model combinations and to exploring the feasibility of abstention-capable triage. For researchers and clinical technology teams asking whether multimodal large language models can serve as foundation models for cystoscopic assessment, it provides concrete reference points for accuracy, calibration, and triage coverage.
The performance gains from prompt engineering were not statistically significant, so their robustness needs observation across more data and settings; the 98.1% triage accuracy corresponds to 62.0% image coverage, and how the remaining images are handled is a question to clarify for practical deployment; in addition, this reading was at summary scope, so stratified results such as performance by imaging modality are not expanded in the text, and those details would affect judgment about the applicable range of the conclusions.
