Large Language Model Data Abstraction Demonstrates Accuracy and Reliability for NSQIP
Synopsis
Using clinical notes from 105 patients in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025), manually de-identified and processed with a customized ChatGPT 4.1 workflow targeting individual variables against a faculty plastic surgeon reference standard, this study evaluated 9,048 data points and found overall abstraction accuracy of 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction, with McNemar and Chi-square p<0.001; the LLM exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables, and the most frequent LLM errors involved prior breast surgical history (29/61) and prepectoral versus subpectoral implant or expander placement.
Interpretation
For NSQIP breast reconstruction variable abstraction, the customized LLM achieved higher overall accuracy than conventional human abstraction, 99.33% (61 errors) versus 98.19% (164 errors), with McNemar and Chi-square p<0.001. NSQIP data collection previously depended on labor-intensive manual chart abstraction; this study directly compared LLM variable-by-variable abstraction with human abstraction under the same reference standard, providing a quantified accuracy comparison in this setting. A comparative evaluation of 105 patients and 9,048 data points, benchmarked against a reference standard established by a faculty plastic surgeon and analyzed with McNemar and Chi-square tests.
LLM performance exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables. This indicates the model's advantage is not uniform across variable categories, and variable type changes the relative performance of model versus human abstraction. Derived from the same 9,048 data points compared by variable category.
The most frequent LLM errors involved prior breast surgical history (29 of 61 errors), followed by prepectoral versus subpectoral implant or expander placement, a variable frequently requiring inference from documentation. It identifies the specific variable types where the model errs, pointing toward targeted improvement of abstraction workflows for variables requiring inference. Based on a categorized count of the 61 model errors by variable.
The authors frame the work as a proof-of-concept validation study, concluding that the findings support LLM-assisted abstraction as a promising approach to improve efficiency, reduce resource requirements, and facilitate broader implementation and expansion of NSQIP, while noting that multicenter validation remains necessary. It translates single-center accuracy results into a judgment about the efficiency and scalability prospects of the NSQIP data collection workflow. The conclusion rests on a validation design within a single pilot program, and the authors explicitly identify multicenter validation as a necessary next step.
Perspective
This work applies to manually de-identified clinical notes from patients enrolled in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025) and to the general and breast reconstruction NSQIP variables targeted by the customized ChatGPT 4.1 workflow, with the reference standard established by a faculty plastic surgeon. Its value is for teams seeking to use model-assisted abstraction to replace or supplement manual abstraction, thereby easing efficiency and cost pressures and reducing patient sampling driven by staffing limits. The authors explicitly state that broader implementation and expansion still require multicenter validation.
A careful reader would still watch how model performance varies with documentation quality for variables requiring inference, such as prepectoral versus subpectoral implant or expander placement; whether the high-frequency error category of prior breast surgical history can be improved through prompt or workflow adjustments; and how the multicenter validation the authors call for would affect accuracy and efficiency conclusions. This is a research letter without figures or variable-level detail, so the specific accuracy values by variable category cannot be further verified.
