Skip to main content
Back to timeline
Scientific ReportsSource publication:

Scaling vision models does not consistently improve localisation-based explanation quality

Synopsis

This study evaluated 11 vision models from the ResNet, DenseNet, and Vision Transformer families (seven trained from scratch, four pretrained) on the Oxford-IIIT Pet and Chest X-ray Pneumothorax datasets with ground-truth segmentation masks, generated explanations with five post-hoc XAI methods (Saliency, GradientSHAP, Integrated Gradients, Feature Permutation, Grad-CAM), and quantified mask alignment using Relevance Rank Accuracy and the proposed Dual-Polarity Precision; it found that increasing depth and parameter count did not improve explanation quality in most statistical comparisons, that smaller models often matched or exceeded deeper variants, that pretraining typically improved predictive performance and increased the dependence of explanations on learned weights without consisten

Source-provided article image: Scaling vision models does not consistently improve localisation-based explanation quality.

Fig. 1

PubMed

Interpretation

Across 60 tests of model scale, the largest model within a family achieved the highest score in only 3 comparisons (5%), smaller models matched or exceeded the largest model in 18 significant comparisons (30%), and 39 of 60 tests (65%) did not reject the null hypothesis of equal explanation quality after Benjamini-Hochberg correction, with effect sizes typically negligible to small. Prior work largely assumed or only sporadically observed the link between scale and explanation fidelity; this study provides a systematic comparison across 11 models, 5 explanation methods, 2 datasets, and 3 random seeds. Based on attributions for 200 randomly sampled test images per seed (600 images across three seeds), using Kruskal-Wallis and Mann-Whitney U tests with Benjamini-Hochberg adjustment and reported Cliff's δ and η² effect sizes.

Introduces Dual-Polarity Precision (DPP), which splits the attribution map into positive and negative parts, computes the precision of positive attributions inside the mask and of negative attributions outside it, and takes their macro-average, thereby penalising positive background mass and negative in-mask mass. Existing mass-based evaluation often maps signed attributions to non-negative values by squaring or rectifying before aggregation, which removes sign information so that negative attributions contribute by magnitude rather than as counter-evidence. Provides explicit formulas for TPA, FPA, TNA, FNA, P_pos, and P_neg, and states that the random baseline is 0.5 for signed attributions and 0 for positive-only methods with a maximum alignment score of 0.5.

Across 40 tests of pretraining, 16 (40%) were significant after correction, with pretrained models scoring higher in only 5 (12.5%), scratch-trained models matching or exceeding pretrained ones in 11 (27.5%), and 24 (60%) not rejecting the null; the pretrained ResNet-50 also showed larger degradation under layer-wise weight randomisation, and Saliency sensitivity to pixel permutation rose from 0.003 to 0.082. Extends the effect of pretraining from predictive performance to localisation quality and dependence on learned parameters, indicating that pretraining changes the link between explanations and weights more than it consistently raises localisation scores. Statistical tests used 200 images per seed (600 across three seeds); sanity checks used 32 images per seed on Oxford-IIIT Pet with layer-wise randomisation and pixel permutation, reporting Cliff's δ.

On the Chest X Pneumothorax dataset, RRA values were near zero and DPP values near the 0.5 baseline across models and attribution methods, while classification AUC ranged from 0.77 to 0.87, indicating that predictive performance and explanation quality can diverge. Uses ground-truth-mask localisation metrics to show directly that high classification performance and near-zero localisation alignment can coexist, complementing evaluations based on accuracy alone. Masks in this dataset cover on average only 1.4% of image area versus 29.7% for Oxford-IIIT Pet, which the authors note makes evaluation more sensitive to small localisation errors.

Perspective

This is a localisation-based evaluation under binary classification, applicable to natural and medical imaging settings that have expert segmentation masks and require auditing whether attributions fall inside regions of interest; it offers RRA and DPP as quantifiable metrics for model selection and auditing. The authors note that future work can extend to multi-class problems, concept-based explanation methods such as TCAV, prototype-based interpretable architectures such as ProtoPNet, and controlled out-of-distribution and adversarial settings to test whether RRA and DPP relate to reliability, and suggest multi-rater datasets such as LIDC-IDRI where soft masks from inter-annotator agreement allow probabilistic assessment.

Mask overlap serves as a proxy for explanation quality, which may be incomplete when clinically relevant context lies outside the annotated region; heatmap-based measures do not cover concept-level or object-level explanations; three seeds provide only a minimal estimate of inter-seed variability; power was estimated at α=0.05 rather than the Benjamini-Hochberg adjusted threshold, which the authors describe as slightly optimistic and indicative only; the CheXpert-pretrained DenseNet-121 did not converge within the 100-epoch budget across all three seeds on Oxford-IIIT Pet, and those comparisons are marked for completeness rather than as primary evidence; greyscale inputs may affect non-medical domains where colour carries signal; and whether RRA and DPP relate to expert decision support still requires studies with radiologists.

Sources