Skip to main content
Back to timeline
JMIR AISource publication:

Retrieve-Then-Verify for Evaluating Evidence Support and Hallucination in Large Language Model-Generated Medical Information

Synopsis

This study evaluated three large language models (GPT-5, OpenAI o3-mini, and GPT-3.5) on risk-of-bias (RoB 2) assessment across 97 randomized controlled trials, retrieving relevant passages from trial reports with Okapi BM25 and then assigning each generated claim an evidence verdict of supported, contradicted, not found, or out of scope with verbatim quotations; binary accuracy of AI-generated judgments was high (90%-98%), yet evidence support rates were only 60%-65% with conservative hallucination rates of 34%-37%, showing that high decision-level accuracy does not guarantee documentary support.

Interpretation

It introduces and applies a retrieve-then-verify pipeline that uses Okapi BM25 to retrieve passages from trial reports and then assigns each model-generated RoB 2 claim a verdict of supported, contradicted, not found, or out of scope with verbatim quotations. Prior evaluations often relied on agreement with human judgments; this work directly compares generated content against source documents and quantifies evidence support and hallucination rates. Run on all 97 randomized controlled trials with complete human RoB 2 annotations and accessible full-text reports, constituting the complete reference set, with no train-validation split.

AI-generated risk-of-bias judgments showed high binary accuracy across domains (90%-98%) but substantially lower exact accuracy (42%-71%), indicating frequent disagreement in severity classification even when directional classification was correct. Reporting exact and binary accuracy together reveals grading differences that agreement metrics alone can obscure. Evaluated with exact and binary accuracy, sensitivity, specificity, F1-score, Youden J, and agreement with human reviewers using Cohen κ and Fleiss κ.

GPT-5 achieved the strongest overall performance, including perfect binary accuracy for overall risk-of-bias conclusions, the highest agreement with human reviewers (quadratic κ up to 0.81), the highest mean evidence support (64.3%), and the lowest strict hallucination rate (35.7%). It provides a relative ranking among the three models while presenting both decision performance and evidence support performance. Based on the same 97-trial reference set and uniform structured RoB 2 output constraints.

Mean top-1 BM25 retrieval scores were similar across models (approximately 30-31), suggesting that differences in hallucination were not primarily attributable to differences in retrieval strength. It separates retrieval quality from hallucination differences, indicating the issue is more likely in generation than in retrieval. The same retrieval algorithm and reference set were used for all three models, making retrieval scores directly comparable.

Perspective

The work addresses digital health practitioners and medical informatics researchers who use large language models for evidence synthesis, guideline development, and clinical knowledge management, in settings that require traceable and auditable medical evidence appraisal; its pipeline constrains outputs to structured RoB 2 signaling questions and domain-level judgments and verifies against randomized controlled trial reports with accessible full text.

Readers may still watch how the pipeline performs on unstructured outputs, other evidence appraisal tools, or non-randomized-trial literature; whether results remain stable when the retrieval algorithm or verification rules change; and how comparable evidence support and hallucination rates are across tasks and model versions. The loaded text is summary-level and does not include figures or full methodological detail, so understanding of specific verdict rules and per-domain results remains limited.

Sources