Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

medRxiv

A Large Language Model for Risk-of-Bias Assessment in Systematic Reviews of Prognosis Studies in Clinical Neurology

This study designed a zero-shot prompted LLM pipeline as a virtual mimic of a human reviewer for the QUIPS framework and applied it to 298 articles from previously published systematic reviews of prognosis studies in neurology across three domains (epilepsy, traumatic brain injury, stroke), finding limited LLM-human agreement (weighted kappa = 0.22, 95% CI 0.12-0.33) while tentatively suggesting, based on a small sample (n=5), that it may not be inferior to human-human agreement (weighted kappa = -0.25, 95% CI -1.04-0.54), with Wilcoxon signed-rank tests statistically significant (p < 0.05) across four bias domains and overall risk scores, and rank-biserial correlations demonstrating human raters' tendency to assign higher risk scores than LLM counterparts.