Public articles linked to the same research event.
medRxiv Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.
Using manually annotated cardiac catheterization reports from three hospitals within Yale New Haven Health (1,412 notes, 596 PCI reports) as a reference standard, this study evaluated three open-weight large language models (Llama 3.3 70B, Meditron-7B, BioMistral-7B) for identifying PCI reports and extracting six complex PCI features, finding that Llama 3.3 70B outperformed the two smaller domain-specific models on most tasks, achieving 100.0% sensitivity, 93.8% specificity, 96.4% accuracy and 95.9% F1 for PCI identification, and among 590 evaluable PCI reports 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value and 83.