Skip to main content
Back to timeline
arXivSource publication:

Tencent Hunyuan and collaborators compared 11 MoE rungs and found encoder-free multimodal loss falls faster with compute, crossing over near 10^22 FLOPs

Synopsis

Using a matched ladder of 11 sparse MoE decoders (1.1B–44B total, 71M–2.4B active non-embedding parameters) with identical data mixture and optimization, the work fits separate text and multimodal scaling laws for encoder-free and encoder-based MLLMs and finds that removing the visual encoder shifts compute-optimal multimodal allocation toward larger models (model allocation exponent rising from about 0.464 to about 0.570) while leaving text nearly unchanged; encoder-free models underperform at small scales but their multimodal loss falls faster with compute, with extrapolation predicting an efficiency crossover on the order of 10^22 FLOPs (point estimate about 3.

AI-generated editorial illustration: How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Interpretation

Removing the visual encoder changes compute-optimal allocation for the multimodal objective: the model allocation exponent rises from about 0.464 for encoder-based to about 0.570 for encoder-free models, with the data exponent falling correspondingly from about 0.536 to about 0.430, while the text objective stays nearly identical (about 0.427 versus about 0.422). Prior encoder-free MLLM work focused mainly on feasibility, alignment, or distillation; this work runs a controlled scaling comparison against an encoder-based baseline on the same decoder ladder and turns the question of whether a larger decoder is needed into a quantified exponent difference. Eleven shared MoE decoder rungs with matched data mixture, optimization, and visual-token granularity; the multimodal allocation gap comes with conditional bootstrap central intervals (about 0.570 encoder-free, about 0.464 encoder-based) and persists under a causal-attention variant.

The two architectures have nearly overlapping loss–compute frontiers on the text objective but diverge on the multimodal objective: encoder-free models need more compute for equal validation loss within the measured range, yet their loss–compute exponent is larger, and extrapolation predicts a crossover on the order of 10^22 FLOPs (point estimate about 3.2×10^21 FLOPs, conditional bootstrap interval about 1.6×10^21 to 1.1×10^22), arriving later under overtraining. Earlier work mostly reported that encoder-free models lag at a given scale; this work places the gap in a scaling-law extrapolation and gives the compute order at which the crossover is predicted, compared against the roughly 10^25 FLOPs pretraining compute of recent flagship models. Loss–compute laws fitted to six IsoFLOP optima, with leave-one-budget-out refits placing the crossover between about 1.6×10^21 and 1.1×10^22 FLOPs; the authors state the projection assumes the fitted laws persist beyond the measured range, the visual encoder stays fixed in size, and irreducible loss is set only by the data distribution.

The decoder takes over visual encoding through vision-specific adaptation: bidirectional attention among visual tokens becomes more beneficial with compute (causal attention saves about 1.5% compute on text but is about 0.5% worse on multimodal on average), visual representations diverge from their inputs at shallower layers while text trajectories stay close across architectures, and expert routing for visual tokens becomes more concentrated (higher MaxVio and a wider band). These internal probes turn the idea that the decoder compensates for the missing encoder from speculation into an observable phenomenon, and localize the effect to visual processing rather than a whole-model routing shift, since text-token routing stays closely matched. In the 8B models, attention to visual tokens at layer 12 rises from about 0.2 to about 0.6 in the encoder-free model, approaching the encoder-based model; layerwise cosine-similarity probes repeat across model scales; expert load imbalance is measured with MaxVio and holds across the four largest model sizes.

Broken down by topic, encoder-free models catch up in a different order: STEM is already close to parity, Charts narrows quickly, and GUI, OCR, and Caption lag considerably longer; on downstream benchmarks encoder-free models still score below encoder-based models at the evaluated scales, but the gap narrows as training compute increases and, at the largest token budget, also tends to narrow as the model grows. Splitting aggregate multimodal loss by topic shows that when parity arrives depends on how strongly a topic relies on pretrained visual representations, rather than following one uniform schedule. Topic-level IsoFLOP profiles reuse existing checkpoints without training new models, with some topic optima near or beyond the ladder edge handled by mild extrapolation; downstream evaluation is 3-shot with no further training and reports averages over perception, document-understanding, and general-VQA benchmarks.

Perspective

The results apply to decoder scaling under a fixed-size pretrained visual encoder, a fixed data mixture, and a fixed MoE topology, and are aimed at research and engineering teams working on multimodal pretraining architecture and compute allocation; the work offers a fitted-within-range, extrapolated-to-about-10^22-FLOPs prediction that can inform when an encoder-free route is worth pursuing and which topics (such as STEM and Charts) approach parity first.

The crossover depends on assumptions that the fitted laws persist beyond the measured range, that the visual encoder size is fixed, and that irreducible loss is set only by the data distribution; the authors describe these as structural assumptions rather than asymptotes identifiable from finite-scale observations. Exponents and crossover points are sensitive to the assumed irreducible loss, and the reported ranges are sensitivity ranges rather than statistical confidence intervals. Some topic-level optima sit near or beyond the ladder edge and involve mild extrapolation. Downstream evaluation still shows encoder-free models behind at the evaluated scales, so the crossover prediction awaits validation at larger scale.

Sources