Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

medRxiv

An LLM-Enabled Pipeline for Natural History Study Information Extraction in Rare Disease Research

This study built a proof-of-concept information extraction pipeline using three open-source LLMs (Athena-v3-AWQ, Gemma3-27B, and Llama-3.1-70B-Instruct) to extract 11 natural history study characteristics from PubMed abstracts labeled "2" in the CZI DRSM corpus (148 gold-standard and 3,547 full-corpus abstracts), finding that all three models exceeded 99% processing success, that Gemma was best overall on expert rating (68.0% of outputs rated "good") and full-corpus runtime (~16 minutes), that Llama scored higher on automated Token F1 (0.874 vs. 0.723), and that Athena performed worst largely because it copied source text verbatim rather than synthesizing it.