Public articles linked to the same research event.
Research Square In this exploratory cross-sectional study, 68 investigator-developed, evidence-mapped English prompts, each containing a false, absolute, unsafe, or contested premise, were submitted once in independent single-turn conversations to ChatGPT (GPT-5), Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, yielding 340 complete responses that five clinicians independently scored, showing that 307 responses (90.3%) were classified safe and 33 (9.7%) unsafe, with safe-response rates from 85.3% to 94.1% and a non-significant omnibus safety comparison (P = 0.203) that was not interpreted as equivalence, while all six graded outcomes differed across models (all P < 0.
In this exploratory cross-sectional study, 68 investigator-developed, evidence-mapped English prompts, each containing a false, absolute, unsafe, or contested premise, were submitted once in independent single-turn conversations to ChatGPT (GPT-5), Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, yielding 340 complete responses that five clinicians independently scored, showing that 307 responses (90.3%) were classified safe and 33 (9.7%) unsafe, with safe-response rates from 85.3% to 94.1% and a non-significant omnibus safety comparison (P = 0.203) that was not interpreted as equivalence, while all six graded outcomes differed across models (all P < 0.
In this exploratory cross-sectional study, 68 investigator-developed, evidence-mapped English prompts, each containing a false, absolute, unsafe, or contested premise, were submitted once in independent single-turn conversations to ChatGPT (GPT-5), Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, yielding 340 complete responses that five clinicians independently scored, showing that 307 responses (90.3%) were classified safe and 33 (9.7%) unsafe, with safe-response rates from 85.3% to 94.1% and a non-significant omnibus safety comparison (P = 0.203) that was not interpreted as equivalence, while all six graded outcomes differed across models (all P < 0.
In this exploratory cross-sectional study, 68 investigator-developed, evidence-mapped English prompts, each containing a false, absolute, unsafe, or contested premise, were submitted once in independent single-turn conversations to ChatGPT (GPT-5), Gemini 2.5 Pro, Microsoft Copilot, DeepSeek-V3.2-Exp, and Doubao-Seed-1.6, yielding 340 complete responses that five clinicians independently scored, showing that 307 responses (90.3%) were classified safe and 33 (9.7%) unsafe, with safe-response rates from 85.3% to 94.1% and a non-significant omnibus safety comparison (P = 0.203) that was not interpreted as equivalence, while all six graded outcomes differed across models (all P < 0.