Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

World journal of urology

Multimodal Large Language Models for Bladder Tumor Detection in Cystoscopy: A Retrospective Benchmarking Study

This retrospective study analyzed 1,754 labeled public cystoscopy images to test Direct, Book-based, and Optimized prompts across GPT-5.2, GPT-5, GPT-5-Mini, and GPT-5-Nano for benign-versus-malignant classification, finding that GPT-5 and GPT-5-Mini with the optimized prompt reached accuracies of 86.7% and 89.2%, and that GPT-5 with the optimized prompt achieved 98.1% accuracy, 94.6% specificity, and 99.1% sensitivity in high-confidence triage at 62.0% image coverage, while prompt engineering improved calibration and triage without statistically significant performance gains.