Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

JMIR AI

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model-Generated Medical Information

This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.