An LLM-Enabled Pipeline for Natural History Study Information Extraction in Rare Disease Research
Synopsis
This study built a proof-of-concept information extraction pipeline using three open-source LLMs (Athena-v3-AWQ, Gemma3-27B, and Llama-3.1-70B-Instruct) to extract 11 natural history study characteristics from PubMed abstracts labeled "2" in the CZI DRSM corpus (148 gold-standard and 3,547 full-corpus abstracts), finding that all three models exceeded 99% processing success, that Gemma was best overall on expert rating (68.0% of outputs rated "good") and full-corpus runtime (~16 minutes), that Llama scored higher on automated Token F1 (0.874 vs. 0.723), and that Athena performed worst largely because it copied source text verbatim rather than synthesizing it.
Interpretation
It established a reusable NHS extraction pipeline: working with a rare disease expert, 10 research questions were translated into 11 precisely defined characteristics (disease name, study purpose, study type, sample size, data collection period, inclusion criteria, exclusion criteria, clinical outcomes, treatments received, study duration, study results), and extraction quality was improved through iterative prompt engineering. Compared with prior rule-based methods, support vector machines, conditional random fields, or BioBERT-style models that require manual features and annotated training data, this pipeline produces structured JSON directly from open-source LLMs without fine-tuning, and adds explicit definitions and format constraints for easily conflated fields such as clinical outcomes and treatments received. The methods are described in detail, including the initial and final prompt contents, an error-pattern analysis using a pediatric primary sclerosing cholangitis abstract, and identical configurations across models (four GPUs, 90% GPU memory utilization, 3,072-token context, temperature 0.1, maximum output 2,048 tokens, vLLM deployment).
On efficiency and completeness, all three models processed abstracts with success rates exceeding 99%; on the full 3,547-abstract corpus Gemma averaged about 16 minutes, Athena about 18 minutes, and Llama about 40 minutes; Gemma and Llama completed extraction on all abstracts, while Athena failed on 39 of 3,547. The results provide a runtime and failure-rate comparison of open-source models of different scales on a real biomedical corpus under the same hardware, and show that the efficiency ranking on the small gold-standard set (148 abstracts, where Athena was fastest) reversed when the corpus was expanded. Based on runtime statistics from 11 consecutive runs (the first measuring loading and initial inference, the following 10 averaged) plus two-level missing-rate tables for the gold-standard set and the full corpus.
On expert rating Gemma was best: among 50 randomly selected abstracts, Gemma had a mean score of 1.40 with 68.0% rated "good", Llama 1.76 with 36.0% "good", and Athena 2.28 with only 10.0% "good"; automated metrics ranked the opposite way, with Llama at Token F1 0.874 and exact-match F1 0.838 versus Gemma at 0.723 and 0.695. This contrast directly surfaces the divergence between automated metrics and expert judgment, and attributes Athena's low scores mainly to verbatim reproduction rather than synthesis, for example returning near-verbatim narrative text for a checkpoint inhibitor toxicity abstract and a recurrent respiratory papillomatosis abstract. Expert review was performed by a rare disease expert who independently scored all three models' outputs for 50 abstracts on a three-point scale (1 = good, 2 = limited, 3 = poor/errors), presented alongside majority-vote automated metrics on the same subset.
Extraction completeness varied markedly by characteristic: disease name, study purpose, study type, and clinical outcomes had very low missing rates across models, whereas exclusion criteria (40.3%–95.2% missing) and data collection period (5.2%–54.8% missing) had the highest; for terminology standardization, GARD mapping was 70.1%–74.2%, RxNorm 62.4%–77.6%, and HPO only 28.0%–38.2%. The results attribute missing-rate differences both to abstract reporting practices and to differences in model interpretation, and suggest that phenotype extraction may require additional normalization strategies to align free-text descriptions with standard ontologies. Based on two-level missing-rate tables for the gold-standard set and full corpus, plus matched/attempted counts for the GARD, HPO, and RxNorm controlled vocabularies.
Perspective
The pipeline is intended for settings where PubMed abstracts are the input and rare disease literature related to natural history studies is the target, and for institutions that want to batch-extract descriptive study characteristics with open-source models on local compute; its design goal is concise, structured fields suited to downstream analysis rather than replacing full-text reading or clinical judgment. The authors note that incorporating full-text articles and clinical trial records and extending to quantitative outcome extraction could further improve completeness and applicability.
A careful reader would still watch: expert review covered only 50 randomly selected abstracts and was performed by a single reviewer, so the stability of ratings under larger samples and multiple reviewers remains to be seen; model parameter counts differ substantially (Llama 70B, Gemma 27B, Athena 13B), and the authors propose future comparisons among models of similar size or efficiency-normalized metrics; exclusion criteria, data collection period, and treatment information are themselves underreported in abstracts, so how missing rates change with full text and clinical trial records is an open question; the low HPO mapping rate suggests phenotype normalization still needs additional strategies; and this is a preprint that has not been peer reviewed.
