Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

World Journal of Methodology

Do bots provide correct and adequate guidance regarding acidity: A blinded comparison rated by patients and physicians

This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.