Mid-training a language model on raw video lifts video benchmarks by 2.9 points and image benchmarks by 5.1 points while preserving text ability
Synopsis
This work asks whether raw web video, with no captions and no text loss, can serve as mid-training data for a pretrained language model: frames are encoded into continuous visual tokens and the model learns to predict the next visual token, Qwen3-1.7B is mid-trained on raw clips from YT-Temporal-1B and then given the same image-text instruction tuning as a model without mid-training, yielding 2.9 points higher on video benchmarks and 5.1 points higher on image benchmarks, with text average rising from 48.0 to 48.9, and caption prediction offering no advantage over visual-token prediction.
Figure 1: Overview of continuous video mid-training. Ordered video frames x 1 : T x_{1:T} are encoded and projected into continuous token representations. Spatial patch grids are serialized with row-delimiter tokens and concatenated into a unified sequence V V . A causal language model processes the continuous visual stream without captions or text supervision, and a prediction head r ω r_{\omega} regresses the subsequent visual representation v i + 1 v_{i+1} from the hidden state h i h_{i} . A stop-gradient operator sg ( ⋅ ) \operatorname{sg}(\cdot) is applied to the regression target so that gradients do not flow through it.
arXivInterpretation
Proposes a mid-training objective that needs no captions or annotations: each frame is mapped by a vision encoder and projector into continuous visual tokens, serialized in row-major order with a learned newline embedding, and the causal language model performs self-supervised feature regression to predict the next visual token. Prior multimodal work either trains unified transformers from scratch or adds latent future-frame prediction heads over curated clips; this work instead continues training an existing pretrained LLM, with no discrete visual codebook and no language supervision. The objective minimizes cosine distance between the predicted representation and the ground-truth visual token, with a stop-gradient on the target so gradients flow only through the causal predictive pathway; under causal self-attention this single objective covers spatial next-patch modeling within frames and temporal prediction across frame boundaries.
Video mid-training improves both video and static image understanding: the average over four video benchmarks rises 2.9 points, including 4.8 points on EgoSchema, and all ten image benchmarks improve, averaging 5.1 points, with ChartQA up 8.2, DocVQA up 6.4, and AI2D up 6.1. Gains transfer from video to static images and concentrate in text-rich and chart tasks, indicating that representations from predictive mid-training over continuous visual tokens are reusable across modalities. The comparison holds the image-text instruction tuning fixed, so the two models differ only in mid-training; training is end-to-end for one epoch (27,300 steps) on 128 NVIDIA H100 GPUs with an effective batch size of 256 clips and a 16k-token context window.
Although mid-training contains no text, text ability is preserved: the average over 14 text benchmarks rises from 48.0 to 48.9, and every knowledge and commonsense score stays within 1.4 points of the model without mid-training. Unlike the common practice of mixing text-only data to limit language loss, this setup uses only image-text instruction data, and the mid-trained model degrades less on GSM8K, MATH500, and BBH (drops of 18.6, 16.5, and 9.5 points versus 23.5, 17.7, and 15.0 for the model without mid-training). Text evaluation spans commonsense and world knowledge, mathematical and multi-step reasoning, and code generation; the authors offer one possible reason, that mid-training has already adapted the language model to visual tokens so instruction tuning needs to change it less.
Predicting captions does not beat predicting visual tokens directly: a next-frame caption target gives similar video and image scores but a 1.5-point lower text average, and a current-frame caption target falls 5.0 points behind visual tokens on image understanding. This indicates video mid-training can remain fully self-supervised, avoiding the expense of frame-by-frame captioning pipelines and the label noise of automated captions. Caption baselines are generated by Qwen3-VL-30B-A3B in consecutive pairs with a progress-aware prompt; when controlling for training budget against the 0.5-epoch visual model (52.37 video, 53.76 image, 48.32 text), visual next-token prediction matches vision gains within 0.3 points while better preserving text capability.
Perspective
The result applies to multimodal pipelines that start from a pretrained causal language model and then undergo image-text instruction tuning: in that setting, mid-training on 16-frame raw clips from YT-Temporal-1B yields video and image gains while holding text performance. Methodologically it adds only a small prediction head during mid-training and requires no discrete codebook or language supervision, so it can extend to large web video collections without annotation. The authors note that whether denser temporal sampling or longer sequence horizons can yield further improvements remains open, and that evaluating these dynamics across larger language models will establish whether uncaptioned video mid-training can serve as a standard precursor to multimodal instruction tuning.
Gains appear within the first 0.3 epochs (19.5B visual tokens) and then vary by less than 0.5 points over the remaining 0.7 epochs, so whether longer training adds anything is unclear. The authors also note that determining whether these gains stem from temporal structure or simply from additional exposure to visual tokens requires future frame-shuffling experiments. On text, both models score below the original language model on GSM8K, MATH500, and BBH, and the mid-trained model's 5.5-point margin on BBH is smaller than the 11.7-point discrepancy between instruction-tuning runs from the same initial state, so the robustness of the text-preservation effect remains to be seen. Reasoning benchmarks also fluctuate by up to 9 points across evaluation intervals, which readers should keep in mind when interpreting single-point scores.
