Skip to main content
Back to timeline
arXivSource publication:

StreamMAE brings MAE to continuous video streams, reaching 74.0 mIoU on Cityscapes with WT++95h and matching same-data i.i.d. pretraining

Synopsis

The authors build WT++, a 95-hour urban walking-tour video dataset, benchmark MoCo v3, DINO and MAE under a strict sliding-window streaming regime without global shuffling or multi-epoch replay, find that contrastive and self-distillation methods lag while MAE is more robust but still falls short of i.i.d. pretraining, use a controlled experiment to attribute the gap to high intra-batch rather than inter-batch similarity, and propose StreamMAE, which keeps the MAE reconstruction objective while adding stream-aware regularization and motion-biased crop selection, outperforming streaming baselines, matching same-data i.i.d. MAE, and improving as the pretraining stream grows from 12 to 95 hours.

AI-generated editorial illustration: I Have a Stream: Making Self-Supervised Learning Work on Continuous Video

Interpretation

Under a strict streaming regime, masked reconstruction suits continuous video better than contrastive learning or self-distillation. Prior work on continuous video typically relies on pretrained initialization, replay buffers, or relaxed from-scratch conditions; this work trains from random initialization with sliding-window batches consumed in temporal order, without global shuffling or multi-epoch replay, and systematically compares three objective families. With ViT-S/16 on WT++12h and full fine-tuning, MAE reaches 77.1 Acc@1 on ImageNet-1K, 61.3 mIoU on Cityscapes and 23.9 mIoU on ADE20K, ahead of DINO (72.7 / 53.2 / 23.9) and MoCo v3 (68.7 / 52.0 / 19.4), with MoCo v3 even below random initialization (71.9).

The gap between streaming and i.i.d. pretraining is driven mainly by high intra-batch similarity rather than inter-batch similarity. The authors consume a once-pre-shuffled ImageNet-1K as a fixed stream, preserving the high inter-batch similarity of sliding windows while keeping batches diverse, thereby decoupling the two factors, a controlled design not previously applied to this attribution question. DINOv2 feature statistics give WT++12h intra-batch similarity μintra=0.665 and inter-batch μinter=1.000, versus 0.004 and 0.989 for pre-shuffled ImageNet-1K; the pre-shuffled stream reaches 77.4 Acc@1, 63.6 mIoU and 27.2 mIoU with ViT-S, essentially matching standard i.i.d. MAE (77.4 / 64.0 / 26.9). On WT++12h, lowering intra-batch similarity from 0.665 to 0.171 improves ViT-B ADE20K and Cityscapes by 6.1 and 10.6 mIoU.

StreamMAE narrows and largely closes the streaming gap while keeping the MAE reconstruction objective unchanged. Rather than changing the loss, the method adapts the input pipeline: color jitter, higher drop path, DataDrop, two-stage cropping, and motion-biased crop selection based on patch-level L1 differences between consecutive frames. With ViT-S/16 on WT++12h it reaches 77.5 Acc@1, 63.8 mIoU and 26.1 mIoU, beating streaming MAE (77.1 / 61.3 / 23.9) and matching same-data i.i.d. MAE (77.0 / 63.5 / 25.9); with ViT-B/16 on WT++12h it reaches 81.1 / 69.0 / 32.7, exceeding same-data i.i.d. MAE on Cityscapes (67.8). Cumulative ablations show regularization and crop selection contribute most, while DataDrop gives modest gains.

Performance scales with pretraining stream duration and model capacity, and transfers beyond walking-tour video. The authors extend WT++ from 12 hours to ordered 25-, 50- and 95-hour streams, and repeat the comparison of streaming MAE, StreamMAE and same-data i.i.d. MAE on three additional domains: HD-EPIC, CROWD and KrishnaCAM. For ViT-B/16 from WT++12h to WT++95h, ImageNet-1K rises from 81.1 to 82.0 Acc@1, Cityscapes from 69.0 to 74.0 mIoU and ADE20K from 32.7 to 36.5 mIoU; on the three extra domains segmentation improves by 1.6–5.8 mIoU and depth RMSE drops by 0.042–0.531, with HD-EPIC and CROWD matching or exceeding same-data i.i.d. MAE on every task.

Perspective

The results apply to a streaming pretraining setting that starts from random initialization, consumes frames in temporal order, and uses neither global shuffling nor multi-epoch replay, which suits embodied agents and edge devices where data arrive sequentially. WT++ comprises 95 hours across 58 public urban walking-tour videos, and the authors plan to release metadata, source links and preprocessing scripts under CC BY 4.0, with 56 of the 58 videos redistributable. Methodologically, apart from motion-biased crop selection, StreamMAE does not explicitly model temporal dynamics; its gains are validated on ViT-S and ViT-B, on 12- to 95-hour streams, and on three additional video domains, making it a baseline and starting point for subsequent streaming pretraining work.

The authors' own open questions include: apart from motion-biased crop selection, StreamMAE does not explicitly model temporal dynamics; fixed-rate frame subsampling may be suboptimal because the rate of visual change varies substantially within a video, so adaptive temporal sampling is worth exploring; and checkpoint averaging is currently used only at evaluation time, leaving long-timescale consolidation inside streaming pretraining unexplored. In addition, DataDrop gives modest gains at ViT-S scale but becomes negligible for ViT-B on longer streams, so its scope of usefulness still needs observation across more settings. The MemoryStoryboard reproduction is a controlled reproduction under the ViT-S/16 protocol, and the authors note that the original configuration uses larger strides, which may leave room for further optimization.

Sources