Skip to main content
Back to timeline
World Journal of MethodologySource publication:

Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians

Synopsis

This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.

AI-generated editorial illustration: Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians.

Interpretation

In physician ratings, Gemini-2.5 and Claude-4 achieved higher mean scores for accuracy, comprehensiveness, and actionability compared with ChatGPT-5 (P < 0.05), while Claude-4 demonstrated the highest empathy scores. Prior concerns about LLMs for gastrointestinal health information centered on accuracy, empathy, actionability, and readability, but there was a lack of direct multi-model comparison on the specific patient topic of "acidity" with blinded gastroenterologist ratings. Three gastroenterologists independently rated responses to 39 frequently asked questions, with statistically significant inter-model differences reported (P < 0.05).

Patient ratings indicated uniformly high comprehensibility across all models, but Gemini-2.5 and Claude-4 responses were perceived as more actionable than those generated by ChatGPT-5. It places the patient perspective alongside physician ratings, complementing prior work that relied mainly on expert or automated evaluation of LLM patient education content. Twenty patients rated empathy, comprehensiveness, actionability, compassion, and usefulness.

Readability analysis showed that ChatGPT-5 produced the most accessible responses, corresponding approximately to a high-school reading level, whereas Gemini-2.5 and Claude-4 generated more linguistically complex content. It reveals a trade-off between readability and physician-rated dimensions: models scoring higher on physician ratings also produced more linguistically complex content. Based on readability index analysis of responses from the three models.

Perspective

This work applies to common patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) and is relevant to clinicians, patient-education designers, and those selecting models for AI-assisted patient education; its findings suggest model selection and hybrid approaches may be considered in gastroenterology patient education, but because results are based on 39 questions rated by three gastroenterologists and 20 patients, applicability to other disease topics or patient groups requires further study.

Readers may wonder about the specific mean scores and distributions for each model across accuracy, comprehensiveness, empathy, and actionability, which are not given in the text; the specific readability indices used are not specified; and the selection criteria for the 39 questions, agreement between patient and physician ratings, and whether the same patterns hold for gastrointestinal topics beyond "acidity" are directions for further observation.

Sources