Public articles linked to the same research event.
Journal of the American College of Surgeons Using clinical notes from 105 patients in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025), manually de-identified and processed with a customized ChatGPT 4.1 workflow targeting individual variables against a faculty plastic surgeon reference standard, this study evaluated 9,048 data points and found overall abstraction accuracy of 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction, with McNemar and Chi-square p<0.001; the LLM exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables, and the most frequent LLM errors involved prior breast surgical history (29/61) and prepectoral versus subpectoral implant or expander placement.
Using clinical notes from 105 patients in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025), manually de-identified and processed with a customized ChatGPT 4.1 workflow targeting individual variables against a faculty plastic surgeon reference standard, this study evaluated 9,048 data points and found overall abstraction accuracy of 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction, with McNemar and Chi-square p<0.001; the LLM exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables, and the most frequent LLM errors involved prior breast surgical history (29/61) and prepectoral versus subpectoral implant or expander placement.
Using clinical notes from 105 patients in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025), manually de-identified and processed with a customized ChatGPT 4.1 workflow targeting individual variables against a faculty plastic surgeon reference standard, this study evaluated 9,048 data points and found overall abstraction accuracy of 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction, with McNemar and Chi-square p<0.001; the LLM exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables, and the most frequent LLM errors involved prior breast surgical history (29/61) and prepectoral versus subpectoral implant or expander placement.
Using clinical notes from 105 patients in the NSQIP Breast Reconstruction pilot program (July 1, 2024–February 28, 2025), manually de-identified and processed with a customized ChatGPT 4.1 workflow targeting individual variables against a faculty plastic surgeon reference standard, this study evaluated 9,048 data points and found overall abstraction accuracy of 99.33% (61 errors) for the LLM versus 98.19% (164 errors) for human abstraction, with McNemar and Chi-square p<0.001; the LLM exceeded human abstraction for operative and postoperative variables but was slightly lower for preoperative variables, and the most frequent LLM errors involved prior breast surgical history (29/61) and prepectoral versus subpectoral implant or expander placement.