Public articles linked to the same research event.
Journal of medical Internet research This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.
This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.