Noise and reverberation interventions shift Alzheimer's predictions across three self-supervised speech backbones
Related research and updatesSynopsis
Using ADReSSo and three large self-supervised speech backbones, the study applies controlled noise and reverberation interventions to participant-speech-only, non-speech, and full-recording audio, and combines layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis to show that controlled acoustic interventions alter Alzheimer's disease predictions across all three backbones, that noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects, and that these effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set, and reverse when the representation-space intervention direction is reversed.
Figure 2: E2&E4, input-space intervention&alignment. (a) AD prediction flip rates across layers for SNR and SRMR interventions; colors denote streams, gradients degradation severity D1–D4, and the dashed line marks ℓ AD ∗ = 12 \ell_{\mathrm{AD}}^{*}=12 . (b) HC → \rightarrow AD and AD → \rightarrow HC flips at ℓ AD ∗ = 12 \ell_{\mathrm{AD}}^{*}=12 , averaged across levels. (c) Cosine alignment with the AD decision direction at ℓ AD ∗ = 12 \ell_{\mathrm{AD}}^{*}=12 ; gray regions show 95% random-direction intervals.
arXivInterpretation
Controlled acoustic interventions alter Alzheimer's disease predictions across all three large self-supervised speech backbones, indicating that acoustic factors are not merely encoded in representations but can systematically change predictions. Prior work largely asked whether acoustic factors can be decoded from self-supervised representations; this work moves to the intervention level and separates decodability from influence on prediction. Controlled noise and reverberation interventions on ADReSSo across three backbones, covering participant-speech-only, non-speech, and full-recording audio, combined with layer-wise linear decoding, input- and representation-space interventions, and geometric alignment analysis.
Noise, despite showing no significant diagnostic-group difference in the original data, produces the strongest intervention effects. This challenges the intuition that a measured acoustic factor without a significant group difference cannot influence the model. From the comparison of noise and reverberation intervention effects, observed consistently across the three backbones.
The intervention effects are systematically structured relative to the classifier's decision direction, replicate on the held-out test set, and reverse when the representation-space intervention direction is reversed. The reversal under direction inversion provides evidence beyond correlational observation, moving the link between acoustic factors and prediction to a direction-controllable level. Representation-space intervention with direction reversal, replication on the held-out test set, and geometric alignment analysis.
High predictive performance and the absence of a significant diagnostic-group difference in a measured acoustic factor are not sufficient for robustness. This shifts robustness assessment from performance metrics and group-level statistical tests toward behavior under intervention. Based on consistent intervention effects across three backbones, held-out test set replication, and direction reversal; the authors argue intervention-based robustness tests should become standard for trustworthy clinical speech models.
Perspective
The work targets speech-based Alzheimer's disease assessment, using the ADReSSo dataset and three large self-supervised speech backbones, with interventions limited to controlled noise and reverberation and audio conditions split into participant-speech-only, non-speech, and full-recording. Within this setting, it suggests to researchers and developers of clinical speech models that robustness evaluation should include intervention tests at both the input and representation levels, and should check whether effects are systematically structured relative to the classifier's decision direction, replicate on held-out data, and reverse with the intervention direction. The authors argue intervention-based robustness tests should become standard for trustworthy clinical speech models, a claim whose scope is speech-based clinical assessment models.
This is presented as an abstract, without specific intervention strengths, effect-size values, layer-wise decoding results, or quantitative geometric alignment metrics, so the actual magnitude of effects and how much they differ across backbones cannot be judged from the available text. Why noise produces the strongest intervention effects despite no significant diagnostic-group difference in the original data remains to be explained mechanistically. How intervention-based robustness testing would be operationalized as a standard procedure, including what thresholds or control conditions are needed, also remains an open question.
