Skip to main content
Back to timeline
Research SquareSource publication:

Clinical quality of large language model answers to questions about pancreatic cancer

Synopsis

This study had eight large language models (ChatGPT-5.5 Instant, Claude Sonnet 4.6, Gemini 3 Flash, Grok 4.20, Microsoft Copilot, Doubao Seed 1.8, DeepSeek-V4, and Qwen 3.6) answer ten patient-oriented pancreatic cancer questions on three consecutive days, with five evaluators blinded to platform identity rating all 240 responses across six dimensions of clinical quality (1,200 ratings), finding that overall performance differed among platforms, with Grok 4.20 and DeepSeek-V4 having the highest mean scores and Doubao Seed 1.8 the lowest, differences being greatest in actionability and specificity, adjacent-day score fluctuation being most pronounced for Doubao Seed 1.8 and Qwen 3.

AI-generated editorial illustration: Clinical quality of large language model answers to questions about pancreatic cancer

Interpretation

The study systematically scored the clinical quality of model answers at a scale of ten patient-oriented pancreatic cancer questions, three consecutive days of submission, eight platforms, five blinded evaluators, six dimensions of clinical quality, 240 responses, and 1,200 ratings. Compared with prior evaluations often focused on a single platform or a single round of questioning, this work combines multiple platforms, multi-day repetition, and blinded multi-dimensional rating so that cross-platform comparison and short-term consistency can be observed within one framework. The evidence comes from structured rating data of 240 responses and 1,200 ratings, with evaluators blinded to platform identity and coverage of six dimensions of clinical quality.

Overall performance differed among platforms, with Grok 4.20 and DeepSeek-V4 having the highest mean scores and Doubao Seed 1.8 the lowest, and the differences were most evident in actionability and specificity. This provides a head-to-head comparison under the same question set showing that clinical quality in patient-facing cancer Q&A is not equivalent across models, and it locates the differences in dimensions such as actionability and specificity that are more relevant to patients' actual decisions. The conclusion is based on score comparisons of the same ten questions across eight platforms, with the magnitude of differences presented through mean scores and cross-dimension contrasts.

Adjacent-day score fluctuation was most pronounced for Doubao Seed 1.8 and Qwen 3.6, but no platform showed a significant systematic day effect after correction for multiple testing. This result treats short-term consistency as an independent object of observation and suggests that day-to-day fluctuation and systematic day effects are two things that need to be distinguished. The evidence comes from adjacent-day score changes obtained through three consecutive days of repeated submission, with the judgment of no significant systematic day effect made after correction for multiple testing.

The evaluated platforms generally provided relevant information, but clinical usefulness and repeatability varied, so LLM-generated pancreatic cancer information should complement rather than replace professional clinical communication. The study moves the evaluation focus from whether information is relevant to whether it is clinically useful and repeatable, and on that basis offers positioning advice for patient use. This judgment rests on rating results across six dimensions of clinical quality and three days of repeated observation, and is a summative conclusion about performance within the evaluated scope.

Perspective

This work applies to the specific setting of patient-oriented pancreatic cancer Q&A, covering eight named platforms, ten questions, three consecutive days, and six dimensions of clinical quality, so its conclusions most directly serve patients, clinical communicators, and evaluation designers who want to understand the clinical quality and short-term consistency of model answers; use in other cancer types, other languages, or other questioning formats would require separate verification.

A careful reader would still watch: whether ten questions and a three-day observation window can represent broader patient questioning and longer-term use; how the specific composition and scoring rules of the six dimensions of clinical quality affect cross-platform differences; how to understand short-term consistency when adjacent-day fluctuation coexists with no significant systematic day effect; and whether conclusions still hold after platform version updates. Because the text read here is of incomplete scope, if figures or supplementary materials contain dimension details and statistical results, the relevant judgments should still defer to the original.

Sources