Public articles linked to the same research event.
JMIR AI This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.
This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.