Generative AI Responses to Patient-Presented Orthodontic Misinformation and Controversial Claims: A Comparative Cross-Sectional Evaluation of Safety, Accuracy, Evidence Concordance, Misinformation Correction, Uncertainty Communication, Transparency, and Actionability
Synopsis
In this exploratory cross-sectional study, 68 investigator-developed, evidence-mapped English prompts, each containing a false, absolute, unsafe, or contested premise, were submitted once in independent single-turn conversations to ChatGPT (GPT-5), Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, yielding 340 complete responses that five clinicians independently scored, showing that 307 responses (90.3%) were classified safe and 33 (9.7%) unsafe, with safe-response rates from 85.3% to 94.1% and a non-significant omnibus safety comparison (P = 0.203) that was not interpreted as equivalence, while all six graded outcomes differed across models (all P < 0.
Interpretation
The study constructed and froze a question-specific reference standard for 68 orthodontic misinformation and controversial claims (evidence freeze 14 August 2026), specifying for each question the misinformation or claim archetype, risk tier, must-include elements, unsafe/must-not content, a critical or high-severity error anchor, evidence mapping, and an overall evidence-certainty anchor, together with a separate operational scoring manual translating that framework into fixed item-level anchors. Prior orthodontic large language model research largely emphasized routine questions, readability, or general answer quality; this study shifts the stimulus to prompts that deliberately preserve false, exaggerated, absolute, unsafe, or contested premises and supplies auditable per-question evidence anchors, making active premise correction a scorable construct. The reference framework was developed by two investigators and checked by a third, with differences resolved within the study team before formal scoring; it was used for evaluator training and adjudication and was not shown to the tested chatbots; the complete question-specific reference standard is in Supplementary Appendix 3 and the item-by-item scoring map in Supplementary Appendix 4.
The study scored binary safety and six graded outcomes (accuracy, evidence concordance, misinformation correction, uncertainty communication, transparency, and actionability) as distinct constructs and reported high inter-rater agreement: Fleiss' kappa = 0.889 (95% CI 0.828–0.936) for safety and ICC(2,1) from 0.864 (transparency) to 0.890 (accuracy) for the six graded outcomes. This design separates avoidance of overtly harmful advice from whether an answer is complete, corrects the premise, and offers a usable next step, preventing a single global rating from concealing omission of a safety-critical element. Five clinicians with 6–26 years of clinical experience completed a calibration exercise on 25 pilot responses not included in the final dataset; scoring files had chatbot names, provider identifiers, interface labels, and logos removed and were presented with coded identifiers in randomized order; all 340 responses were scored independently before any consensus discussion and the independent ratings were locked for agreement analysis.
Of 340 responses, 307 (90.3%) were classified safe and 33 (9.7%) unsafe; safe-response rates were 94.1% (64/68) for ChatGPT, 92.7% (63/68) for Gemini, 91.2% (62/68) for Copilot, 88.2% (60/68) for DeepSeek, and 85.3% (58/68) for Doubao, with a non-significant omnibus safety comparison (Cochran's Q = 5.949, df = 4, P = 0.203; effect size 0.022, 95% CI 0.009–0.080) and no significant pairwise McNemar comparison after Benjamini-Hochberg adjustment (adjusted P = 0.653–1.000). The study states explicitly that this non-significant result does not establish equivalence and that no equivalence margin was prespecified, so safety-rate differences should be read as not detected rather than shown to be absent. Safety was binary, and an unsafe classification required a specific plausible pathway to harmful behavior, delayed assessment, unsafe self-treatment, or preventable injury; factual imprecision, incomplete explanation, unsupported certainty, or poor communication alone was not sufficient for an unsafe classification.
All six graded outcomes differed across models (all P < 0.001): mean accuracy was 4.09 ± 0.73 for ChatGPT, 3.81 ± 0.78 for Gemini, 3.66 ± 0.78 for Copilot, 3.41 ± 0.70 for DeepSeek, and 3.03 ± 0.75 for Doubao (Kendall's W = 0.382); median evidence concordance was 75.00 for both ChatGPT and Gemini (interquartile ranges 75.00–87.50 and 62.50–87.50) and 62.50 for Copilot, DeepSeek, and Doubao; median misinformation correction was 80.00 for ChatGPT and Gemini, 70.00 for Copilot and DeepSeek, and 60.00 for Doubao; median uncertainty communication was 75.00 for ChatGPT, Gemini, and Copilot, 62.50 for DeepSeek, and 50.00 for Doubao; median transparency was 75.00 for ChatGPT and Gemini, 62.50 for Copilot and DeepSeek, and 50.00 for Doubao; median actionability was 100.00 for ChatGPT, 75.00 for Gemini and Copilot, and 66.67 for DeepSeek and Doubao. The results support reporting model rankings by domain rather than collapsing them into a single claim of universal superiority: ChatGPT led mean accuracy, but the median evidence-concordance result was tied with Gemini, and the two had different interquartile ranges. The matched design had the same 68 questions answered by all five systems; overall differences used Friedman tests, post hoc comparisons used paired Wilcoxon signed-rank tests, and Kendall's W served as the omnibus effect size; after Benjamini-Hochberg adjustment, 58 of 60 pairwise comparisons were significant, with exceptions being Gemini versus Copilot for accuracy (adjusted P = 0.098) and Copilot versus DeepSeek for actionability (adjusted P = 0.091).
Perspective
The study defines the scope of its conclusions: the object of evaluation is the provider-hosted consumer chatbot products available during the recorded query sessions rather than fixed underlying model architectures; the tested systems were ChatGPT, Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, and for ChatGPT, Copilot, and Doubao the exact backend routing or deployed snapshot could not be independently verified from retained interface metadata. The study evaluated generated response text rather than diagnostic performance, treatment effectiveness, patient comprehension, or clinical outcomes; the prompts were standardized research scenarios rather than a patient-validated questionnaire and were not prevalence-weighted. Safety coding targeted prespecified harm pathways, so the high overall safe-response proportion should be read as avoidance of a prespecified harm pathway rather than confirmation that answers were complete, personalized, or clinically effective. Actionability denotes the presence of a clear and safe next step and is not proof that a patient would understand or follow it. No clinician, search-engine, or guideline comparator was included, and patient comprehension, trust, decisions, adherence, outcomes, and downstream harm were not measured; bias and fairness were not evaluated as separate outcomes.
Several open questions remain for careful readers: the products were time-stamped and may change behavior through updates, routing, retrieval functions, or safety-policy changes without a new public model name, so the findings should not be assumed to remain stable; each English prompt was submitted once, and multi-turn, multilingual, and repeated-generation performance were not tested, with repeated-query reproducibility not assessed; the investigator-developed question set was not patient-validated or prevalence-weighted, and some topic domains were small; the structured response archive preserves answer text and plain-text representations of source tables but does not reproduce embedded images or other non-text interface elements, so the evaluation applies to the archived textual content; complete reviewer blinding was not possible because style could reveal a model; the study-specific scales were not externally validated patient-reported instruments; unsafe events were uncommon and no equivalence margin was prespecified, with topic analyses descriptive. Future studies could use prospectively collected patient questions, multilingual prompts, repeated runs, browsing comparisons, and multi-turn challenges, together with blinded external panels, clinician and evidence-synthesis comparators, and preregistered equivalence designs, and could link content scores to comprehension, calibrated trust, care-seeking, and treatment decisions.
