Skip to main content
Back to timeline
UrogynecologySource publication:

ChatGPT-generated surgical decision aids for pelvic floor disorders had high patient understandability, but about 30% of topics fell short on accuracy and overall readability was high

Synopsis

In this cross-sectional study, six urogynecologists developed six questions comparing two treatment options for common urogynecological conditions and entered them into ChatGPT to create decision aid tools; patients and physicians then rated understandability with a Patient Education Materials Assessment Tool, physicians rated reliability with a modified DISCERN instrument and accuracy on a 5-point Likert scale, and readability was assessed with the Flesch-Kincaid Reading Ease score, showing high patient and physician understandability for all tools, fair reliability with an average mDISCERN score of 26, accuracy below 4 (unfavorable) on 2 of the tools, and a high reading level required overall.

Source-provided article image: Evaluation of AI-Generated Surgical Decision Aids for Pelvic Floor Disorders.
PubMed

Interpretation

ChatGPT-generated decision aid tools for urogynecology surgical counseling received high understandability from both patients and physicians. Prior studies had begun evaluating whether artificial intelligence can create patient education materials; this study moves the evaluation to decision aid tools for specific surgical decisions and includes both patient and physician raters. Cross-sectional study with six questions developed by six urogynecologists; patients and physicians used a Patient Education Materials Assessment Tool, and analysis was descriptive.

Reliability was only fair, with an average mDISCERN score of 26. Adds a physician-rated reliability dimension beyond understandability, so the evaluation is not limited to readability alone. Physicians rated reliability using a modified DISCERN instrument, with an average score of 26 reported and descriptive analysis.

Two of the six decision aid tools scored below 4 (unfavorable) on accuracy, and approximately 30% of the topics require improvements to accuracy. Quantifies accuracy as a separate dimension, showing that understandability does not guarantee accuracy and giving a concrete target for improvement. Physicians rated accuracy on a 5-point Likert scale; two tools scored below 4, and the conclusion states that approximately 30% of topics require accuracy improvements.

The reading level required for all tools was high overall, leaving room to improve readability. Introduces the Flesch-Kincaid Reading Ease score alongside understandability ratings, revealing that high understandability and high reading difficulty can coexist. Readability was evaluated using the Flesch-Kincaid Reading Ease score, and the conclusion states that all tools could benefit from improved readability.

Perspective

This study characterizes ChatGPT-generated urogynecology surgical decision aid tools across four dimensions: understandability, reliability, accuracy, and readability. It applies to counseling scenarios in which urogynecologists pose questions comparing two treatment options for common urogynecological conditions. For clinical teams exploring AI-assisted shared decision making, the result suggests patient acceptance may not be the main barrier, and next steps can focus on improving reliability and readability and validating across broader conditions and populations.

Readers should still watch: the descriptive analysis provides no statistical inference, the scale of six questions and six urogynecologists is limited, the number and composition of patient and physician raters are not given in the text, and which topics correspond to the two tools scoring below 4 on accuracy is not specified. In addition, the specific ChatGPT version and prompting approach are not described in the text, so whether results would hold under different models or prompts remains an open question.

Sources