Skip to main content
Back to timeline
medRxivSource publication:

A Large Language Model for Risk-of-Bias Assessment in Systematic Reviews of Prognosis Studies in Clinical Neurology

Synopsis

This study designed a zero-shot prompted LLM pipeline as a virtual mimic of a human reviewer for the QUIPS framework and applied it to 298 articles from previously published systematic reviews of prognosis studies in neurology across three domains (epilepsy, traumatic brain injury, stroke), finding limited LLM-human agreement (weighted kappa = 0.22, 95% CI 0.12-0.33) while tentatively suggesting, based on a small sample (n=5), that it may not be inferior to human-human agreement (weighted kappa = -0.25, 95% CI -1.04-0.54), with Wilcoxon signed-rank tests statistically significant (p < 0.05) across four bias domains and overall risk scores, and rank-biserial correlations demonstrating human raters' tendency to assign higher risk scores than LLM counterparts.

AI-generated editorial illustration: A large language model for risk-of-bias assessment in systematic reviews of prognosis studies in clinical neurology

Interpretation

The study demonstrates the feasibility of a tailored, prompt-engineered LLM pipeline for automating risk-of-bias assessments with the QUIPS tool. Prior applications of LLMs in systematic review automation focused on other steps; this study is the first to design a zero-shot prompted LLM pipeline specifically as a virtual mimic of a human reviewer under the QUIPS framework, focused on a single discipline of neurology prognosis research. Based on 298 articles from 15 systematic reviews across three domains (epilepsy, traumatic brain injury, stroke), demonstrating pipeline feasibility, though specific technical details of pipeline construction are not reported.

LLM-human agreement on risk-of-bias assessments was limited, with a weighted kappa of 0.22 (95% CI 0.12-0.33). Provides a quantitative estimate of LLM-human agreement under the QUIPS framework, filling a prior gap in such data for risk-of-bias assessment in prognosis studies. Based on assessments of 298 articles with a relatively narrow confidence interval, though the agreement level itself is low.

Preliminary data based on a small sample (n=5) tentatively suggest that LLM-human agreement may not be inferior to human-human agreement (weighted kappa = -0.25, 95% CI -1.04-0.54). First comparison of LLM-human agreement alongside human-human agreement under the QUIPS framework, providing an initial reference for the relative performance of LLM assessment. Very small sample (n=5) with an extremely wide confidence interval; the authors explicitly describe this as 'tentatively suggest,' indicating limited evidence strength.

Wilcoxon signed-rank tests were statistically significant (p < 0.05) across four bias domains and for overall risk scores, and rank-biserial correlations demonstrated human raters' tendency to assign higher risk scores than LLM counterparts. Reveals the direction of systematic differences in risk score distributions between LLM and human assessors, namely that humans tend to assign higher risk scores. Based on paired comparisons across 298 articles with statistically significant tests, though the specific implications of the effect direction require interpretation within the QUIPS framework's scoring criteria.

Perspective

This study provides feasibility evidence for LLM-automated risk-of-bias assessment under the QUIPS framework, applicable to systematic reviews of prognosis studies in neurology (epilepsy, traumatic brain injury, stroke). The authors note that with targeted methodological refinements—including standardization of QUIPS implementation and validation against expert ratings—automated risk-of-bias assessment may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond. Potential beneficiaries include systematic review authors, methodologists, and evidence-based medicine practitioners.

LLM-human agreement is limited (weighted kappa = 0.22), and the human-human agreement comparison is based on only n=5 with an extremely wide confidence interval (-1.04 to 0.54), leaving the conclusion of whether LLM assessment is non-inferior to human assessment uncertain. The systematic difference in which human raters tend to assign higher risk scores warrants further investigation into its sources and implications. Additionally, this study is a summary-level rapid parse lacking flow diagrams, tables, and detailed score distributions across QUIPS domains; the absence of this information affects a complete understanding of assessment details. Standardization of QUIPS implementation and validation against expert ratings are key future directions proposed by the authors.

Sources