REViT-v2 cuts roto-reflection equivariant attention cost from quadratic to linear via windowed group-convolutional self-attention, beating rotation/flip-augmented ViT-S and RE-ResNet on ImageNet-1K
Synopsis
The authors propose REViT-v2, which restricts group-convolutional self-attention to local windows (wG-CSA) and builds a multi-stage equivariant feature pyramid from group-convolutional tokenization and downsampling, reducing the spatial complexity of roto-reflection equivariant attention from quadratic to linear in image area; on ImageNet-1K the equivariant REViT-v2-S reaches 79.27% and 80.9% Top-1, above rotation/flip-augmented ViT-S (72.08%) and RE-ResNet (77.37%), and on Rotated MNIST it keeps about 98.26% accuracy with smaller model size, FLOPs, latency, and peak memory.
Interpretation
Introduces Windowed Group Convolutional Self-Attention (wG-CSA), which computes attention inside non-overlapping local windows while preserving complete group-representation fields in each attention head, and introduces no absolute positional embedding, with spatial structure preserved by the equivariant convolutional stem and convolutional Q/K/V projections. Prior G-CSA evaluated globally keeps a spatial attention matrix quadratic in input size, a cost especially important for equivariant models because regular group representations typically allocate several orientation channels per feature field; wG-CSA makes that cost linear in image area. The paper gives a complexity derivation: global attention is about O(n²), while windowed attention with window size w has n/w² windows each requiring w⁴ pairwise comparisons, giving total cost linear in image area for fixed w; Appendix A.1 proves equivariance under orthogonal group representations and the condition that the transformation maps complete windows to complete windows, satisfied by square windows under rotations and axis-aligned reflections when square feature-map dimensions are divisible by the window size.
Builds a multi-stage equivariant feature pyramid: the input image is treated as a field of trivial representations, lifted into regular group representations by a group-equivariant convolutional stem, with two strided group-equivariant convolutions reducing resolution by a factor of four; after the blocks in each of the first three stages, a strided group convolution reduces spatial resolution while increasing the number of representation fields. Transfers hierarchical and multiscale ideas validated in efficient ViTs (Swin's local windows and pyramid, CvT's convolutional token embedding and projections, MViT's progressive resolution reduction with increasing channels) into the group-equivariant attention framework, so that at deeper stages a fixed window corresponds to a progressively larger region of the original image, modeling high-level interactions without constructing a global high-resolution attention matrix. Supported by design description and complexity argument, plus a 'Global w/downsampling' control on Rotated MNIST to distinguish the effect of early downsampling from that of windowed attention.
On ImageNet-1K classification, equivariant REViT-v2 outperforms both baselines: REViT-v2-S reaches 79.27%/94.45% (18 M params) and 80.9%/95.1% (47 M params), above rotation/flip-augmented ViT-S (72.08%/89.54%, 22 M) and RE-ResNet (77.37%/93.74%, 11 M); REViT-v2-T reaches 72.58%/90.88% with 5 M params. Prior equivariant attention formulations did not scale beyond low-resolution images; this work demonstrates group-equivariant ViTs can scale to millions of parameters and ImageNet-scale images. Results come from ImageNet-1K classification experiments; training used four RTX 4090 GPUs with per-GPU batch size 128 (effective batch 512), 300 epochs of AdamW, initial learning rate 3e-4, weight decay 0.05, 20 epochs of linear warmup followed by cosine decay, and window size 7.
On Rotated MNIST (200 epochs, batch size 128), windowed attention reduces forward FLOPs from 1.69 GFLOPs to 80.1 MFLOPs, latency to 0.61 relative to global, and peak training memory from 429.7 MB to 26.47 MB, while keeping comparable accuracy (98.26% vs 98.23%). The comparison also reports an intermediate 'Global w/downsampling' condition (333.6 MFLOPs, latency 0.75, memory 50.98 MB, accuracy 98.28%), separating the efficiency gains of early downsampling from those of windowed attention. Efficiency and accuracy numbers come from a controlled table under the same Rotated MNIST training setup, with comparable parameter counts (97.7 K, 97.7 K, 102.9 K).
Perspective
The work targets large-scale image recognition with equivariant ViTs as backbone feature extractors; the paper states that with suitable interfaces it also applies to dense prediction tasks, but the main text reports ImageNet-1K classification and a complexity comparison on Rotated MNIST. The equivariance proof applies when group representations are orthogonal and the transformation maps complete windows to complete windows, which the paper notes holds for square windows under rotations and axis-aligned reflections when square feature-map dimensions are divisible by the window size. The quantified efficiency gains come from Rotated MNIST, and ImageNet-scale efficiency numbers are not given in the main text. Code and pretrained weights are available at https://github.com/kc-ml2/revit for reproduction and follow-up research.
The equivariance error analysis uses randomly initialized models with no trained checkpoint loaded, so it reflects equivariance induced by network design rather than a property learned through optimization; for the D4 group the paper reports errors accumulating from the lifting layer and attributes this to interpolation artifacts from rotations at angles such as 45 degrees, an author attribution. Equivariance error was measured on 1024 ImageNet validation images processed with resize to 256, center cropping, and ImageNet mean and standard deviation normalization. The windowed-attention efficiency numbers come from a small-scale Rotated MNIST comparison, so their behavior at ImageNet scale and larger model sizes remains for readers to judge against their own settings.
