Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

Journal of medical Internet research

Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.