Skip to main content
Back to timeline
Studies in health technology and informaticsSource publication:

Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs: A Technical Case Report

Synopsis

This technical case report develops and evaluates a system that combines local large language models with rule-based validation against HUGO Gene Nomenclature to extract and structure mutational eligibility criteria from clinical trial records (identifying mutated genes, distinguishing inclusion from exclusion criteria, and assigning them to individual study arms); applied to 4,918 clinical trials it produced structured representations for 1,010 studies, and expert review of 42 trials showed 88.1% of studies correctly annotated with 80% precision at the level of individual eligibility criteria, with failures mainly due to hallucinated genes and misinterpreted abbreviations.

AI-generated editorial illustration: Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs - A Technical Case Report.

Interpretation

It presents and implements a system that extracts mutational eligibility criteria from unstructured trial registry text, identifying mutated genes, distinguishing inclusion from exclusion criteria, and assigning criteria to individual study arms. Relative to prior practice of manual reading or simple rule handling of registry text, the work decomposes extraction into gene identification, inclusion/exclusion separation, and arm assignment as operational steps, with an end-to-end implementation. The paper reports system design, integration, and evaluation, and the system was actually run on 4,918 clinical trials, constituting implementation-and-evaluation evidence at the level of a technical case report.

It uses large language models for context-aware extraction and rule-based validation against HUGO Gene Nomenclature to address ambiguous abbreviations and context-dependent meaning. Combining LLM contextual understanding with rule validation against an authoritative gene nomenclature targets the specific difficulty of abbreviation ambiguity in biomedical text. The method description clearly states the division of labor between the two components, but the abstract provides no ablation or controlled comparison to quantify the rule-validation contribution separately.

The system was integrated into the Community Annotated Trial Search (CATS) platform and produced structured output at real scale: 1,010 of 4,918 trials received structured representations of genetic eligibility criteria. Unlike extraction studies that remain offline experiments, this work demonstrates feasibility of deployment within an operational platform and reports its coverage scale. The coverage scale comes from actual system operation; the specific fields of the structured representation and downstream uses are not elaborated in the abstract.

Expert review of 42 trials showed 88.1% of studies correctly annotated and 80% precision at the level of individual eligibility criteria, with failures mainly attributed to hallucinated genes and misinterpreted abbreviations. It provides human-review-level estimates of accuracy and precision and explicitly names the error types, pointing to directions for subsequent precision improvement. Evidence comes from expert review of 42 trials, a small-scale manual evaluation; the precision figure is computed at the individual-criterion level, a different basis from the study-level 88.1% figure.

Perspective

The results apply to the setting of extracting genetic mutational eligibility criteria from clinical trial registry records, suited to predominantly English registry text where mutation requirements can be validated against HUGO Gene Nomenclature, for example patient-trial matching and clinical decision support in precision oncology; their value lies in providing structured eligibility representations for platforms such as CATS, enabling downstream automated processing.

A careful reader would still watch: the expert review covers only 42 trials, and the 88.1% and 80% figures rest on different bases (study level versus individual-criterion level), so stability at larger sample sizes remains to be observed; hallucinated genes and misinterpreted abbreviations are the main error sources, and how dataset design and error mitigation strategies can reduce their impact remains an open question; moreover, the abstract does not elaborate the specific fields of the structured representation, downstream uses, or local LLM configuration details, which affect judgments about reproducibility and transfer.

Sources