In a single-center cohort of 232 papillary thyroid carcinoma patients, an SVM model predicted lymph node metastasis with validation AUC 0.849, and removing BRAF changed AUC by only 0.007
Synopsis
This retrospective study of 232 patients who underwent thyroid surgery at Tongling People's Hospital used LASSO to select six features from 32 candidates (BRAF mutation status, tumor size, capsular invasion, extrathyroidal invasion, multifocality, and TSH), compared seven machine learning algorithms, found the Support Vector Machine best in the validation cohort (AUC 0.849, 95% CI 0.756–0.934; accuracy 0.768), identified TSH and tumor size as the top SHAP contributors, and showed in an ablation analysis that removing BRAF lowered validation AUC from 0.849 to 0.843 (P = 0.671).
FIGURE 1 Flowchart of patient inclusion and exclusion. A total of 1,151 patients undergoing thyroid surgery were initially screened. After excluding 422 patients with non-malignant tumors or unclear pathological diagnoses, 729 patients remained. Among these, 468 patients were excluded because BRAFV600E testing was unavailable, leaving 261 patients with available BRAF results. A further 29 patients were excluded because of incomplete clinical or biochemical data, resulting in a final analytic cohort of 232 patients. No patients were excluded because of age <18 years.
· Page 3Interpretation
In a postoperative cohort of 232 papillary thyroid carcinoma patients, LASSO retained six predictive features at the one-standard-error criterion in the training cohort: BRAF mutation status, tumor size, capsular invasion, extrathyroidal invasion, multifocality, and TSH, which the authors describe as selected predictive features rather than independent prognostic or causal factors. Compared with models relying predominantly on clinical and ultrasonographic characteristics, this study placed a molecular marker alongside biochemical variables in one feature-selection pipeline, and it states explicitly that structured ultrasound descriptors and Hashimoto's thyroiditis were not candidate variables. Single-center retrospective cohort with 232 patients in the final analysis (163 training, 69 internal validation); LASSO was performed only in the training cohort, and the validation cohort was not involved in feature selection or hyperparameter tuning.
Among the seven algorithms, SVM had the highest validation discrimination, with AUC 0.849 (95% CI 0.756–0.934) and accuracy 0.768, versus training AUC 0.912 (95% CI 0.864–0.950) and accuracy 0.834. On the same data, Random Forest fell from training AUC 0.998 to validation 0.697 and XGBoost from 0.976 to 0.788, whereas SVM showed a smaller training-to-validation decline, which is why the authors selected the final model on validation performance rather than training performance. The validation cohort contained 69 patients; 95% CIs for AUC were estimated with 1,000 bootstrap resamples, and calibration curves, Brier score, and decision curve analysis were used to support the choice of SVM.
The ablation analysis showed a limited incremental contribution from BRAFV600E: the full six-variable SVM had validation AUC 0.849 versus 0.843 for the reduced five-variable SVM without BRAF, an absolute difference of 0.007 that was not significant by paired DeLong test (P = 0.671); at a 0.50 threshold the two models had accuracies of 0.768 and 0.797, sensitivities of 0.860 and 0.860, and specificities of 0.615 and 0.692. Rather than treating the association between BRAF and lymph node metastasis as equivalent to incremental predictive value in a multivariable model, the study re-tuned the reduced model within the same training/validation split and compared the two models pairwise. A within-cohort ablation comparison in which the reduced model was re-tuned only in the training cohort, with the AUC difference assessed by paired DeLong test.
SHAP analysis provided both global and instance-level interpretation, with the feature importance ranking showing TSH and tumor size as the most influential variables driving model output, and waterfall and force plots showing how each feature shifts an individual patient's predicted probability from baseline. It adds an interpretability layer to a relatively opaque SVM, so the six-variable model's risk output can be decomposed case by case rather than summarized only by an overall AUC. Game-theory-based SHAP attribution of an already trained model; this is an explanatory analysis and does not establish causal effects.
Perspective
The model is intended for patients with papillary thyroid carcinoma who have already undergone thyroid surgery and for whom BRAFV600E testing and the relevant clinical and biochemical data are available; it estimates individualized postoperative lymph node metastasis risk and is positioned as an adjunct to postoperative pathological assessment rather than a replacement for histopathological diagnosis or guideline-based management. The authors suggest that when pathological nodal assessment is limited, or when observed nodal findings are discordant with other clinicopathological features, a high model-estimated risk may serve as an additional signal prompting closer review of pathological specimens, targeted postoperative cervical imaging, or multidisciplinary reassessment; a low model-estimated risk should not be used to exclude occult nodal disease. The model is not intended as a stand-alone basis for decisions about radioactive iodine therapy, TSH suppression, or reoperation.
The loaded text is incomplete: Figures 1 through 5 and Supplementary Tables 1 through 3 are not included in the loaded material, so the specific shapes of the calibration curves, decision curves, and SHAP plots, as well as the full hyperparameter search spaces in the supplementary tables, cannot be checked here. Open questions the authors raise include the single-center retrospective design and the sample of 232 patients, the outcome distribution of 62.5% lymph node metastasis versus 37.5% non-metastasis, non-uniform extent of lymph node dissection that may cause under-detection of occult metastases and outcome misclassification, possible selection bias from requiring BRAF testing (lymph node metastasis prevalence 62.5% versus 60.3% in those without testing, P = 0.567), the absence of standardized preoperative ultrasound features and Hashimoto's thyroiditis variables, and the fact that several predictors come from postoperative pathology, which limits purely preoperative use. Whether the model remains robust and generalizable in larger prospective multicenter cohorts, and whether integrating it into clinical workflows improves decision-making or patient outcomes, remain open questions.
