Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports Using Large Language Models
Synopsis
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Interpretation
Across 1,412 catheterization notes (596 of which were PCI reports), Llama 3.3 70B achieved 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, maintaining high accuracy at all three hospitals (approximately 97.1%, 100.0% and 95.5%). A prior study on PCI procedural variables used a smaller number of annotated PCI notes and non-open-source models; this study uses a larger sample, open-weight models deployable locally in a secure environment, and three hospitals with differing note structures. Performance was measured against a reference standard created by four interventional cardiology fellows and reviewed by a board-certified interventional cardiologist, with sensitivity, specificity, positive predictive value, negative predictive value, accuracy and F1 computed on 1,386 evaluable reports and 95% Wilson confidence intervals reported.
Among 590 evaluable PCI reports, Llama 3.3 70B achieved 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value, 83.9% accuracy and 72.5% F1 for complex PCI classification, with the highest accuracy at each hospital (approximately 75.6%, 77.7% and 88.2%). The study targets complex PCI as defined by Giustino et al., a procedural phenotype requiring domain judgment rather than surface-level entity recognition, and reports criterion-level binary classification performance alongside overall classification. Complex PCI status was positive when at least one of six criteria was present, negative only when all six criteria were available and none was present, and otherwise indeterminate; this rule was applied separately to the reference standard and model-extracted variables, and classification performance was evaluated only when both statuses could be determined.
Extraction accuracy varied markedly by variable: Llama 3.3 70B reached high exact-match accuracy for stents implanted (96.8%), bifurcation PCI with two stents (96.5%), total stent length (93.8%) and vessels treated (90.1%), and lower accuracy for lesions treated (80.7%) and chronic total occlusion (85.1%). The study stratifies variables by documentation characteristics and reasoning demands, indicating that explicitly stated, numerically defined variables are more amenable to direct extraction while variables requiring cross-section integration and clinical definition judgment show lower accuracy. Exact-match accuracy was calculated as the proportion of evaluable reports in which the model output matched the reference value, with 95% Wilson confidence intervals; criterion-level sensitivity, specificity, positive predictive value, negative predictive value, accuracy and F1 were also reported.
The two smaller medical-domain models behaved very differently: Meditron-7B had 87.2% sensitivity but only 7.8% specificity for PCI identification, while BioMistral-7B had 1.2% sensitivity and 99.4% specificity; for complex PCI classification BioMistral-7B had only 51 evaluable reports, so its estimates should be interpreted cautiously. This comparison suggests that, for this task, model capacity and general language reasoning may matter more than domain-specific pretraining alone, and that smaller models can achieve high apparent accuracy on rare criteria by predominantly classifying reports as criterion negative. All three models used the same prompt template and were used for inference only without fine-tuning, run locally on an NVIDIA A100; Meditron-7B and BioMistral-7B used 4-bit quantization and deterministic decoding, with full confusion-matrix counts and metrics reported.
Perspective
The results apply to retrospective cardiac catheterization reports from three hospitals within a single health system, in settings where complex PCI is characterized by the six criteria defined by Giustino et al. for research and quality measurement, and support inference with open-weight models in a local secure computing environment. The study indicates that explicitly documented, numerically defined variables such as stent number and stent length are best suited to direct extraction, whereas variables requiring cross-section integration and clinical definition judgment, such as lesion count, bifurcation PCI and chronic total occlusion, are better addressed by hybrid pipelines combining rule-based logic or post-processing constraints with model reasoning.
A careful reader would still watch for external validation across broader health systems and differing documentation practices; prospective validation and performance within human-in-the-loop workflows; how performance changes as document length and information density increase or when the task requires identifying missing information; and how small extraction errors propagate into risk modeling or comparative effectiveness research. In addition, this is a preprint that has not been certified by peer review, raw data are not publicly distributed due to privacy and IRB constraints with only a code link provided, and full criterion-level metrics and confusion matrices reside in supplementary materials, so interpretation of individual criteria is limited without those supplements.
