SemanTok puts semantics in early video tokens: a 201M autoregressive model matches a 3.4x larger VideoFlexTok
Synopsis
SemanTok is a flexible-length video tokenizer that feeds frozen DINOv2 features into its encoder and adds lightweight heads that reconstruct those features from each retained token prefix alone, letting a 201M autoregressive model match or beat a VideoFlexTok model 3.4 times its size and improving semantic alignment at every noise level, including pure noise.
Interpretation
SemanTok fuses frozen DINOv2 features into the encoder input and adds dense and class cross-attention heads that reconstruct DINO features from the retained token prefix alone, so every prefix carries the teacher's semantics. Earlier flexible tokenizers such as VideoFlexTok apply a REPA loss only on early decoder hidden states, a target the decoder can partly meet from its noised latent; SemanTok's heads never see the noised latent, so only the tokens can lower their losses. The paper compares against VideoFlexTok at the same codebook, sequence length, and AR recipe on Kinetics-600 and uCO3D, across seven AR sizes (49M-2.29B) and multiple token budgets.
Semantic alignment and fidelity improve at every AR size: on Kinetics-600 gFVD drops 11-24%, gFID 5-17%, and class accuracy rises 25-61%; on uCO3D the gains are 2-13%, 5-8%, and 22-30%. Gains are largest for the smallest AR models, indicating that semantic supervision moves content decisions into the prefix and eases error compounding when capacity is scarce. Each Kinetics-600 cell uses 2,048 generated clips and each uCO3D split 2,560, with paired bootstrap intervals; at a fixed budget an 85M SemanTok model beats VideoFlexTok of every size in gFID.
SemanTok prefixes are cheaper to predict: at 201M the per-token cross-entropy is lower than VideoFlexTok's, and the first 4 tokens nearly match the class accuracy of VideoFlexTok's first 32, with pixel detail deferred to later tokens. This is a prefix-level compression-generation trade-off: semantics early, detail later, so a small model can settle what the clip shows first. Cross-entropy is measured on 1,024 validation clips and token-repeat statistics on 60k training clips, showing the saving is not from repetition.
From pure noise, SemanTok's decoder-REPA readout has higher semantic alignment at every noise level, and it keeps semantic alignment on out-of-distribution classes. This indicates the semantics come from the token prefix rather than the noised latent, and extends to classes never seen in training. The noise sweep uses 256 validation clips with the same noise seed; OOD results use the uCO3D ID/OOD split, where a 201M AR model raises generated class accuracy on OOD classes.
Perspective
The result targets autoregressive video generation and world models that use flexible-length, coarse-to-fine video tokenizers: at inference the application picks a token budget, a small AR model settles what the clip contains, and larger models fix appearance details. Validation covers class-conditioned Kinetics-600 and text-conditioned uCO3D with an AR size ladder from 49M to 2.29B, so it can inform engineering decisions about picking a model for a compute budget. The authors also list future directions: optimizing predictability directly, organizing early prefixes with video-native or language-aligned teachers, and using such prefixes as compact states for action-conditioned world models.
The paper states that at low token budgets it prioritizes semantic alignment over reconstructing colors and appearance details, and that with one token per frame it struggles to serve every objective at once; uCO3D overfits after 66B tokens while the Kinetics-600 tokenizer was still improving at 131B tokens. The appendix noise probe also cannot separate the three token-side changes from one another, so the relative contribution of each remains an open question. Many table values in the homepage evidence bundle were lost during capture, so specific numbers should be checked against the original.
