Skip to main content
Back to timeline
Journal of medical Internet researchSource publication:

Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study

Synopsis

This retrospective study submitted standardized summaries of 107 single-center hepatopancreatobiliary (HPB) multidisciplinary team (MDT) cases four times each to four large language models (GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5), finding that response stability differed significantly across models (Gemini 3 Pro lowest discordance, mean 12.8%, Fleiss κ=0.737; GPT-4o highest, mean 30.2%, Fleiss κ=0.430), that concordance with MDT decisions ranged from 48.6% to 72.9% using the initial response and from 66.3% to 74.5% using the modal response, and that class-wise F1 was consistently lower for surgery (0.400–0.520) than for chemotherapy (0.621–0.836), indicating that concordance alone is insufficient for evaluating LLMs as clinical decision support tools.

AI-generated editorial illustration: Large Language Models in Multidisciplinary Decision-Making for Hepatopancreatobiliary Oncology: Retrospective Comparative Feasibility Study.

Interpretation

The study systematically quantified the repeat stability of LLM treatment recommendations under identical clinical input, finding significant differences across models (P=.01). Prior evaluations of LLMs for clinical decision support largely focused on concordance with expert decisions and rarely tested whether the same input yields the same answer; this study repeated queries across four separate sessions and introduced stability metrics that do not privilege any single reference response, including discordance rate, mean pairwise agreement, and Fleiss κ. Based on 107 consecutive MDT cases, four models, and four repeated queries per case, with a statistically significant stability difference (P=.01) and reported means and standard deviations of discordance rates per model.

Concordance between LLMs and MDT decisions was moderate and shifted with the evaluation definition (initial versus modal response), with the best-performing model differing between the two definitions. The study reports concordance under both initial-response and modal-response definitions alongside Cohen κ and class-wise F1 scores, showing that a single concordance number can mask changes in the relative ranking of models. Concordance ranged from 48.6% to 72.9% for the initial response and from 66.3% to 74.5% for the modal response, with class-wise F1 scores of 0.400–0.520 for surgery and 0.621–0.836 for chemotherapy.

Complete discordance (all four responses differing from one another) occurred in 17 of 107 cases (15.9%) and in none of the 31 anatomically unresectable cases (Fisher exact test, P=.003). The study further localized the clinical situations in which instability concentrates, rather than reporting only an overall stability figure. Fisher exact test on 107 cases (P=.003), with the number and proportion of complete-discordance cases reported.

Recurrent or on-treatment disease (adjusted OR 5.40, 95% CI 1.65–17.68; P=.005), pancreatic tumor location (adjusted OR 7.37, 95% CI 2.03–26.78; P=.002), and low MDT agreement level (adjusted OR 10.33, 95% CI 1.54–69.38; P=.016) were independently associated with complete discordance. The study identifies case-level factors associated with unstable model responses, offering clues about which situations may warrant additional human review. Multivariable-adjusted odds ratios with 95% confidence intervals, all statistically significant.

Perspective

The scope defined by this study is the single-center HPB oncology MDT setting: 107 cases consecutively discussed between September 1, 2024, and August 31, 2025, with standardized case summaries submitted through consumer web interfaces to GPT-4o, GPT-5.2, Gemini 3 Pro, and Claude Sonnet 4.5, each queried four times per case. Within this setting, the results support viewing LLMs as a reasoning-support layer in MDT-like decision environments rather than as a replacement for MDT decisions; for researchers wishing to assess model stability and for clinical teams considering such tools, the study offers a reusable approach to quantifying stability (discordance rate, mean pairwise agreement, Fleiss κ) and clues about case features that may warrant closer review.

A careful reader would still watch: whether the stability metrics retain the same ordering under more repetitions or different session settings; what the improvement in concordance under the modal-response definition implies; whether the consistently lower surgery F1 scores reflect the model's grasp of surgical indications or the definition of MDT options themselves; and whether the associations between complete discordance and recurrent or on-treatment disease, pancreatic tumor location, and low MDT agreement level hold in larger, multicenter settings. In addition, the text available here is summary-level content without figures or supplementary material, so per-case response distributions, confidence interval details, and sensitivity analyses cannot be further checked; these are open questions to keep in mind when reading the full article.

Sources