Under a region-held-out protocol, LBP+SVM identifies hydrogen-charging signatures in 316L stainless steel SEM images with 0.79 balanced accuracy
Synopsis
For SEM micrographs of 316L stainless steel, this work proposes a Leave-One-Region-Out region-held-out cross-validation protocol over 14 spatial regions (8 AR, 6 H2; 31 images) and compares six feature-classifier combinations built on LBP, GLCM, self-supervised convolutional embeddings, and a CNN; the simplest approach, LBP+SVM, performed best with balanced accuracy 0.79, H2 recall 0.69, and H2 precision 0.82, outperforming every deep-learning and combined-feature model, while a group-level permutation test (500 permutations sampled from the 3,003 possible region-to-label assignments) yielded p = 0.008 and Grad-CAM maps from a CNN tended to concentrate on localized surface and grain-boundary features.
Figure 1: Pooled confusion matrix for the LBP+SVM model under Leave-One-Region-Out cross-validation, aggregated across all 14 folds (31 images in total: 18 AR, 13 H 2 ). Diagonal cells are correct classifications; off-diagonal cells are misclassifications. The model reaches a specificity of 0.89 for AR and a recall of 0.69 for H 2 , with 4 false negatives (H 2 classified as AR) and 2 false positives (AR classified as H 2 ).
arXivInterpretation
It proposes and validates a region-held-out (Leave-One-Region-Out) evaluation protocol for SEM micrographs that classifies as-received (AR) versus hydrogen-charged (H2) states in 316L stainless steel. Prior machine-learning models for this characterization often used image-level splits, which leak information between training and test sets when several images come from the same specimen region; this work instead holds out 14 spatial regions (8 AR, 6 H2; 31 images), aligning the evaluation unit with the specimen region. The protocol is executed as LORO cross-validation over 14 regions and 31 images, accompanied by a group-level permutation test (500 permutations sampled from the 3,003 possible region-to-label assignments, p = 0.008), indicating the result cannot be explained by a chance alignment of the region structure.
Among six feature-classifier combinations, the simplest texture approach, LBP+SVM, achieved the best performance, with balanced accuracy 0.79, H2 recall 0.69, and H2 precision 0.82. This outperformed every deep-learning and combined-feature model built on self-supervised convolutional embeddings, GLCM, and combined features, indicating that hand-crafted texture descriptors can still recover a hydrogen-charging signature when samples are few. The comparison covers six combinations including LBP, GLCM, self-supervised convolutional embeddings pretrained on 143 unlabeled SEM images, and a CNN, all evaluated under the same region-held-out protocol.
Grad-CAM maps from the CNN tended to concentrate on localized surface and grain-boundary features. This offers a qualitative correspondence between where the model attends and where hydrogen-induced morphological changes are known to occur, complementing an evaluation otherwise dominated by aggregate classification metrics. The Grad-CAM maps come from a CNN trained on the full dataset and constitute qualitative visualization evidence rather than independent quantitative validation.
Perspective
The protocol targets settings where SEM micrographs are used to distinguish AR from H2 states, and where several images come from the same specimen region so that regions must be held out; the authors explicitly state the same protocol can be extended to larger HE detection studies in other alloy systems. For researchers who wish to reproduce or adopt this evaluation workflow, its value lies in aligning the evaluation unit with the specimen region and in producing a testable statistical conclusion under small-sample conditions.
The current evidence rests on a small sample of 14 regions and 31 images, and the H2 recall of 0.69 for LBP+SVM means missed detections remain possible; the Grad-CAM observation comes from a CNN trained on the full dataset and is qualitative visualization, not directly matched to quantitative results under the held-out protocol. The loaded text is summary-level and does not include figures, tables, or hyperparameter details, so specifics such as feature settings, preprocessing, and threshold selection still need to be confirmed in the original article.
