Public articles linked to the same research event.
World Journal of Methodology This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.
This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.
This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.
This study submitted 39 frequently asked patient questions about "acidity" (heartburn/dyspepsia/gastroesophageal reflux disease) to ChatGPT-5, Gemini-2.5, and Claude-4, had responses independently rated by three gastroenterologists for accuracy, comprehensiveness, empathy, and actionability and by 20 patients for empathy, comprehensiveness, actionability, compassion, and usefulness, and analyzed readability indices, finding significant inter-model differences across multiple physician-rated domains, with Gemini-2.5 and Claude-4 scoring higher than ChatGPT-5 for accuracy, comprehensiveness, and actionability, Claude-4 showing the highest empathy scores, uniformly high patient-rated comprehensibility across all models, patient ratings of Gemini-2.