A radiomics model automatically separated thermograms of 98 CRPS patients from 56 healthy controls with AUC 0.93, outperforming three clinicians
Synopsis
Using 178 thermograms from 98 patients with complex regional pain syndrome (CRPS) and 837 thermograms from 56 healthy controls, the study used a U-Net deep learning model to segment extremities automatically (mean Dice similarity coefficient 0.99), extracted 564 radiomics features per extremity per thermogram (intensity, shape, texture), and built a classification model with the WORC automated machine learning framework under 20x random-split cross-validation; the model distinguished CRPS from healthy control thermograms with a mean AUC of 0.93 (95% CI 0.90-0.97), statistically significantly better than three clinicians' 0.82, 0.79, and 0.69 (P < 0.001).
Schematic overview of the automatic thermography quantification model development. The thermograms (1) are first processed through a convolutional neural network (2) to segment the extremities from the background. Processing steps include feature extraction (3) and the creation of a machine learning decision model (5), using an ensemble of the best 100 workflows from 1,000 candidate workflows (4), where the workflows are different combinations of the different analysis steps (eg, the classifier used). Adapted from Reference 41 : the scans have been adopted to this study and the convolutional neural network has been added.
Interpretation
The study built a fully automatic CRPS thermogram classification pipeline with a mean AUC of 0.93 (95% CI 0.90-0.97), balanced classification rate 0.84, sensitivity 0.75, and specificity 0.92. Earlier CRPS thermography work relied largely on small manually selected regions of interest and the single measure of left-right temperature asymmetry; this study instead segmented the whole extremity automatically and extracted 564 radiomics features spanning intensity, shape, and texture. Based on 178 thermograms from 98 patients with CRPS and 249 thermograms from 56 healthy controls (427 thermograms in the classification dataset), evaluated with 20x random-split cross-validation, with all model optimization confined to the training set via an internal 5x random-split cross-validation.
The model's AUC was significantly higher than visual scoring by three clinicians (0.82, 0.79, and 0.69; P < 0.0001), and among thermograms the clinicians misclassified, the model correctly classified 20 CRPS thermograms and all control thermograms. Direct comparison between automated thermographic analysis and clinician reading was previously lacking; this study used one randomly selected thermogram per subject, scored by three diagnosis-blinded clinicians on a four-point scale that was then binarized. The three clinicians had 4, 3, and 5 years of experience; the model was additionally evaluated in leave-one-out cross-validation and compared with a DeLong test; all three clinicians concordantly correctly classified 34 of 98 CRPS thermograms (35%) and 43 of 56 control thermograms (77%), and all three misclassified the same 29 CRPS thermograms (30%) and 3 control thermograms (5%).
A single feature group performed close to the all-feature model: intensity only AUC 0.93, texture only 0.91, shape only 0.89, whereas conventional left-right mean temperature difference reached only AUC 0.79. This suggests CRPS thermographic information extends beyond temperature asymmetry to spatial texture patterns; in univariate testing 672 of 1128 features (59%) differed significantly between groups, comprising 646/1016 texture, 18/26 intensity, and 8/86 shape features. All feature-group models were evaluated under the same 20x random-split cross-validation; the temperature asymmetry metric was computed on the entire dataset for simplicity.
Model performance did not depend on age or body part: the age-matched subset (268 thermograms from 56 controls and 48 patients) had AUC 0.93, hands only 0.93, feet only 0.96, and automatic segmentation matched manual segmentation (0.93 vs 0.94). These control models were built to check for age bias, to test whether hands and feet need separate models, and to verify that automatic segmentation can replace time-consuming manual segmentation. Patients and controls differed significantly in age (median 48 vs 33 years, P = 0.001), making the age-matched subset analysis pertinent; the segmentation model achieved a mean DSC of 0.99 (95% CI 0.99-0.99) in cross-validation.
Perspective
The result applies to distinguishing unilateral distal upper- or lower-extremity CRPS patients from healthy controls, using a Thermacam SC5000 camera with emissivity set at 0.99 under a standardized acquisition protocol, with model output restricted to one of two classes, CRPS or control. For clinical readers this means thermograms could gain an automatic, quantitative, observer-independent reading reference, particularly for cases with lack of clinical consensus or subtle temperature abnormalities that are hard to see by eye; for researchers, the pipeline (U-Net segmentation plus WORC radiomics classification) is a starting point for extension to other regions such as knees or lower legs and to treatment-monitoring settings.
The model performs only a two-class comparison between CRPS patients and healthy controls, and the text states that thermograms outside this context would be classified into one of the two categories, so it cannot yet separate CRPS from other post-traumatic or neuropathic pain conditions; the authors also note the absence of independent external validation, and that thermography is sensitive to hardware, acquisition protocols, and environmental conditions, leaving generalizability unknown. In addition, 86% of patient thermograms showed one or more protocol deviations versus 21% of control thermograms, and patients and controls differed significantly in age; these factors were addressed through segmentation, an age-matched subset, and simulated postural-deviation controls, but their residual influence remains something a careful reader would watch. CRPS was modeled as a single diagnostic entity, so whether the model can stratify by pathophysiological mechanism, subtype, or treatment response, and how predictions relate to underlying physiology, remain open questions.
