Skip to main content
Back to timeline
Journal of medical systemsSource publication:

Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction & Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes

Synopsis

This report develops CLASS, a Python-based modular large language model pipeline that runs within a secure institutional environment and combines expert-curated concept lists, a task-specific prompt suite, and a schema-constrained output format to extract structured data from clinical notes, and evaluates it exploratorily on a single-center retrospective corpus of pediatric esophageal airway treatment surgery (EATS) operative notes: observed concordance with surgeon adjudication on the 20 longest notes (3,960 note-procedure pairs) was high (F1 0.9967), while CLASS proposed 28 candidate procedure variants or additions, 18 (64.

AI-generated editorial illustration: Exploratory Implementation and Feasibility Report of CLASS (Clinical LLM Abstraction & Structuring System), A Large Language Model Pipeline for Extracting Unstructured Data From Clinical Notes.

Interpretation

The report presents and describes CLASS, a modular pipeline combining expert-curated concept lists, a task-specific prompt suite, and a schema-constrained output format, exporting results to a dashboard for expert review and analysis. Relative to manual abstraction of notes, the work embeds a large language model within a secure institutional environment to form a configurable, reviewable extraction process. Evidence comes from a descriptive report of system architecture and components, at the level of implementation and feasibility rather than a controlled experiment.

CLASS classifies predefined concepts and flags potential variants or novel concepts for expert consideration, proposing 28 candidate procedure variants or additions on EATS operative notes, 18 (64.3%) of which were judged clinically useful. This extends the pipeline's output beyond filling an existing concept list toward suggestions for expanding a specialized concept list. Evidence consists of exploratory counts and surgeon judgment on a single-center retrospective corpus from one service.

On the 20 longest notes and 3,960 note-procedure pairs, observed concordance between CLASS outputs and surgeon adjudication was high (F1 0.9967). This provides a quantitative observation of how closely model extraction approached expert judgment on this specific task. Evidence is a concordance metric from an exploratory evaluation, restricted to the longest notes and adjudicated by a surgeon at a single center.

The surgeon identified 12 additional procedures that CLASS did not surface, indicating room for expert involvement in concept-list expansion and human review. This result points to where expert input remains part of the workflow. Evidence comes from comparison against surgeon adjudication within the same exploratory evaluation, a single-center observation.

Perspective

The work is aimed at clinical and informatics teams working within a secure institutional environment, configured with expert-curated concept lists and a task-specific prompt suite, who need to extract structured information from unstructured clinical notes; its modular design allows configuration for different abstraction tasks and can process large note volumes within standard institutional infrastructure. The reported results are positioned in a single-center, single-service (pediatric esophageal airway treatment surgery) retrospective operative-note setting, described by the authors as a feasibility-level exploration, with broader validation noted as needed to assess generalizability to other tasks and settings.

A careful reader would still watch how concordance and candidate-concept discovery change on longer or shorter notes, in other specialties, and at other institutions; how the expert-review dashboard is used in real workflows and what burden it adds; how concept lists are maintained and versioned as candidate variants are added; and what conditions are needed to move from single-center exploratory observation to broader validation. This report is a summary-level text without figure or table detail, so further understanding of the specific prompt suite contents, schema-constraint details, and evaluation procedure remains an open question.

Sources