AI-Assisted Data Extraction for Systematic Reviews in Education: Empirical LLM Accuracy and the Human-in-the-Loop Tool AIDE
Synopsis
Through a pilot study and a main study, this work used LLMs including Claude 2.1, ChatPDF, GPT-4, Gemini 1.5 Flash, Gemini 1.5 Pro, and Mistral Large 2 to extract explicit and derived variables from 112 studies in a published education review and compared them with human coding, finding higher agreement for explicitly stated data but markedly lower agreement and generally low Cohen's Kappa for categorizing data into predefined categories, and on that basis proposed and developed the open-source human-in-the-loop (HIL) data extraction tool AIDE that enforces per-item human validation.
Figure 1. LLM results for data extracted that were explicitly stated in the papers.
· Page 7Interpretation
In the context of education systematic reviews, LLMs extract explicitly stated data more accurately than derived data that must be categorized according to predefined categories. Most prior LLM data-extraction evidence came from medical and medically adjacent fields, with limited education evidence; this study replicates and extends that contrast in education using a larger corpus of 112 studies. In the main study, Gemini 1.5 Pro reached 83.23% exact-match agreement (k = 0.40) and 85.02% accurate-match agreement (k = 0.41) on explicit data, dropping to 65.48% (k = 0.24) and 70.24% (k = 0.29) on derived data, with a sample of 112 studies and 24 variables each.
Although precision was generally high, agreement with human coders (Cohen's Kappa) did not reach levels typically acceptable in research synthesis, so LLMs should not serve as fully automated primary extraction tools. The study reports not only accuracy but also precision, recall, F1, and Kappa, and distinguishes 'Exact Match' from 'Accurate Match', revealing insufficient consistency masked by high precision. In the main study precision mostly ranged from 0.89 to 0.95 while recall was lower at 0.58 to 0.80, and Kappa ranged from 0.12 to 0.41; confusion matrices showed models more often extracted inaccurate information than stated information was 'not reported' when it actually was.
Based on these empirical results, the authors proposed and implemented a human-in-the-loop (HIL) workflow and the open-source tool AIDE that enforces per-item human validation. The work translates the empirical conclusion that LLMs cannot be fully trusted into usable software: first the R-based AIDE (run locally, supporting free APIs and local models via Ollama), then a web version of AIDE using structured JSON parsing instead of regex, storing API keys in sessionstorage, and deliberately omitting a 'record all' function. The tool is free and open source with GitHub repositories and OSF data and scripts; the authors report that informal use and colleague feedback suggest the workflow saves time, but they explicitly note this is anecdotal and lacks empirical support.
Perspective
The results apply to the data extraction step of systematic reviews and meta-analyses in education, especially for teams using free or locally available LLMs who are willing to validate every data point by hand; the tool targets researchers who want to reduce text-parsing burden while retaining human judgment, and it recommends reporting in manuscripts which software was used, which LLM was used, and to what extent humans validated the extracted data.
The authors note that time savings, reduced workload, and usability benefits of AIDE currently rest on anecdotal evidence and need quantification; prompt, parameter, and model differences make cross-model and cross-context generalization cautious; there is no existing benchmark for data extraction in education systematic reviews; the performance of newer frontier LLMs remains an open question; and although this was a full-text reading, figures appear as images, so specific confusion-matrix values can only be understood from the prose descriptions.
