XACT learns sparse attribution masks over invertible time-frequency transform coefficients, highlighting fewer spurious features than baselines on synthetic data and yielding sparse, structured explanations on two real-world datasets
Synopsis
The authors propose XACT, a general framework that learns sparse attribution masks over coefficients from arbitrary invertible time-frequency transforms (STFT, continuous wavelet transform, discrete wavelet transform), and extend the virtual inspection layer from the STFT to both wavelet transforms so that LRP can generate explanations in these representations; on a synthetic dataset XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines, and across two real-world datasets it produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria.
Figure 1: Overview of the proposed method. A mask, 𝑴 \boldsymbol{M} , of the same dimensions as the time-frequency representation, g ( 𝒙 ) g(\boldsymbol{x}) , is initialised randomly. An explanation is then the product of a learning problem tied to the mask using an objective function that enforces sparsity, smoothness and consistency between the original prediction of the model and the prediction of the masked input, g − 1 ( g ( 𝒙 ) ∘ 𝑴 ) g^{-1}(g(\boldsymbol{x})\circ\boldsymbol{M}) .
arXivInterpretation
XACT expands the explanation search space from the raw time domain or a fixed transform domain to coefficients of arbitrary invertible time-frequency transforms, learning sparse attribution masks there. Existing attribution methods typically operate either in the time domain or in a fixed transform domain, limiting their ability to capture salient information across representations; XACT learns masks directly on transform coefficients, covering STFT, the continuous wavelet transform, and the discrete wavelet transform. The abstract states the framework was evaluated on the STFT, the continuous wavelet transform, and the discrete wavelet transform, which supports generality at the method level.
The authors extend the virtual inspection layer approach from the STFT to both the continuous and discrete wavelet transforms, enabling LRP to generate explanations in these representations. The virtual inspection layer had been used with the STFT; generalizing it to two wavelet transforms means LRP-based explanations are no longer confined to a single transform domain. The abstract states this extension explicitly but does not provide quantitative comparison details for it.
On a synthetic dataset, XACT produces precise explanations and is less prone to highlighting spurious features than the tested baselines. This provides a controlled setting with known salient feature locations to directly test explanation precision and the tendency to highlight spurious features. The abstract reports a qualitative comparison on the synthetic dataset without giving specific metric values.
Across two real-world datasets, XACT produces sparse and structured explanations, although no method performs best across all quantitative evaluation criteria. This indicates that explanation quality on real data is multi-dimensional, with sparsity and structure coexisting with different evaluation criteria rather than one method dominating all of them. The abstract reports results on two real-world datasets and explicitly states that no method is best across all criteria.
Perspective
This work targets researchers and practitioners who need to explain deep time-series classifiers, in settings where discriminative information lies in latent frequency or time-frequency structure. The method assumes an arbitrary invertible time-frequency transform; the abstract's evaluation covers the STFT, the continuous wavelet transform, and the discrete wavelet transform, tested on one synthetic dataset and two real-world datasets. For readers who want to compare time-domain and multiple transform-domain explanations within one framework, this setup provides a reusable starting point.
The abstract does not list the specific names of the synthetic and two real-world datasets, sample sizes, or metric values, nor does it specify the composition of the baselines, so the size of the gaps between methods cannot be judged from the abstract. The abstract notes that no method performs best across all quantitative evaluation criteria but does not say on which criteria XACT leads or lags. The actual explanation quality after extending the virtual inspection layer to wavelet transforms would also require the main text's comparisons to assess.
