Skip to main content
Back to timeline
Physics in Medicine and BiologySource publication:

GPT-assisted radiomic modeling for predicting pathological complete response to neoadjuvant chemoimmunotherapy in head and neck squamous cell carcinoma

Synopsis

In a training cohort (n = 186), a validation cohort (n = 116), and a prospective multicenter validation cohort (n = 269), this study extracted radiomic and supervised deep learning features from pretreatment T2-weighted MRI and compared manually developed with GPT-assisted modeling workflows, finding that fused features achieved the highest AUC for both manually developed and GPT-assisted logistic regression models (0.759 versus 0.763 ± 0.003 in the training cohort, 0.714 versus 0.741 ± 0.008 in the validation cohort, and 0.700 versus 0.706 ± 0.

AI-generated editorial illustration: GPT-assisted radiomic modeling for predicting pathological complete response to neoadjuvant chemoimmunotherapy in head and neck squamous cell carcinoma.

Interpretation

The study evaluated GPT as an assistant in radiomics modeling, where GPT was not used as the classifier but assisted with preprocessing, feature selection, hyperparameter optimization, code generation and execution, model training, and probability output. It positions GPT as an automation assistant for the modeling workflow rather than as the predictive model itself, differing from the common approach of using AI directly as a classifier. Systematically implemented across three cohorts (training cohort n = 186, validation cohort n = 116, prospective multicenter validation cohort n = 269), with each GPT-assisted workflow independently repeated five times.

Fused radiomic and deep learning features achieved the highest AUC for both manually developed and GPT-assisted logistic regression models. In the setting of predicting pCR to neoadjuvant chemoimmunotherapy in head and neck squamous cell carcinoma, it shows the advantage of fused features over single feature types and provides a side-by-side comparison of manual and GPT-assisted workflows. AUCs were 0.759 versus 0.763 ± 0.003 in the training cohort, 0.714 versus 0.741 ± 0.008 in the validation cohort, and 0.700 versus 0.706 ± 0.004 in the prospective multicenter validation cohort.

GPT-assisted workflows achieved performance broadly comparable to manually developed models across five classifiers, including support vector machine, naïve Bayes, random forest, and XGBoost. It extends the feasibility of GPT-assisted modeling from a single model to multiple classifiers, providing cross-model comparative evidence. Compared across five classifiers using fused features, with GPT-assisted workflows repeated five times and AUC standard deviations ranging from 0.001 to 0.019, indicating limited variability.

Perspective

The study applies to patients with head and neck squamous cell carcinoma receiving neoadjuvant chemoimmunotherapy, using radiomic and supervised deep learning features extracted from pretreatment T2-weighted MRI and GPT-assisted modeling under expert supervision; its conclusions concern the setting of pathological complete response prediction, where GPT-assisted workflows achieved performance broadly comparable to manually developed models and may reduce technical workload and facilitate integrated model development.

The current text is summary-level and does not include figures or full methodological details; readers may watch for the stability of GPT-assisted workflows across different centers and MRI acquisition parameters, and for how the difference between AUCs of 0.700 versus 0.706 in the prospective multicenter validation cohort performs in broader populations, which can be framed as open questions for further validation.

Sources