LAST reuses one set of Transformer blocks to refine only the class token, beating a twelve-layer sequential model on AudioSet by 2.1% with 49.4% fewer parameters
Related research and updatesSynopsis
LAST first processes all tokens, then reuses the same blocks to refine only the class token over fixed audio features, making later passes inexpensive; on AudioSet, ten-pass LAST reaches 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput, while increasing the pass count from two to ten improves accuracy with only 1.2% added computation and shows improved robustness to temporal masking and other auditory augmentations plus better generalization on music, environmental, and event sound classification.
Figure 1: LAST outperformed a deeper transformer at lower cost: AudioSet Computation per clip (x-axis) v/s AudioSet mAP (y-axis) Sequential AST uses D ∈ { 3 , 6 , 9 , 12 } D\in\{3,6,9,12\} layers; the recurrent models use D = 6 D=6 shared blocks and R ∈ { 2 , 4 , 6 , 8 , 10 } R\in\{2,4,6,8,10\} passes.
arXivInterpretation
LAST concentrates additional processing on integrating already-computed features: it first processes all tokens, then reuses the same blocks to refine only the class token, making later passes inexpensive. Relative to sequential transformers that add layers to improve recognition, LAST trades depth for repeated refinement of the class token rather than new parameters. The abstract gives the mechanism and an AudioSet comparison, but no ablation detail or statistical tests.
On AudioSet, ten-pass LAST achieves 0.345 mean average precision, exceeding a twelve-layer sequential transformer by 2.1% relative. This accuracy gain comes together with 49.4% fewer parameters, 42% fewer multiply-accumulate operations, and 9.8% higher measured throughput. The abstract reports accuracy and efficiency numbers on a single dataset, without multiple runs or confidence intervals.
Across separately trained models, increasing the pass count from two to ten improves accuracy while adding only 1.2% computation. This indicates the benefit comes from repeated refinement itself rather than factors shared with the sequential model. The abstract supports the trend with comparisons between separately trained models, without per-pass curves or variance.
Further evaluations show LAST is more robust to temporal masking and various other auditory augmentations, and generalizes better on music, environmental, and event sound classification. Robustness and cross-category generalization are examined alongside sequential depth. The abstract summarizes these evaluations qualitatively, without specific augmentation types, magnitudes, or values.
Perspective
The work targets audio recognition, especially sound event and audio classification on AudioSet; its design intent is to repeatedly refine the class token over fixed audio features at low cost, so it suits settings that want deeper effective processing without more parameters, such as deployments limited by compute or throughput. The robustness and generalization conclusions in the abstract concern temporal masking and other auditory augmentations and music, environmental, and event sound classification tasks.
The abstract does not describe the looped block structure, how class-token refinement is implemented, training configuration, or evaluation protocol, nor the specific augmentation types and magnitudes used in robustness tests. The two-to-ten-pass comparison is based on separately trained models, and per-pass trends and variance are not presented in the abstract. Because only the abstract was available here, figures and body details are not included, and the scope of the reported numbers and conclusions should be checked against the original text.
