Multimodal Flow unifies language and vision as continuous hyperchunks under one flow-matching backbone, with MF-1 scoring 0.821 on GenEval and 83.44 on DPG-Bench at 150B tokens
Synopsis
The work introduces Multimodal Flow, which encodes text blocks and images as ordered continuous hyperchunks and models both modalities in their own embedding spaces with a shared chunk-causal Flow Matching backbone, instantiated as MF-1: continued pretraining improves multimodal modeling across 0.6B, 1.2B, and 1.6B scales, and the 1.6B model reaches 0.821 on GenEval and 83.44 on DPG-Bench with only 150B pretraining tokens and an average of 75.3 across VQAv2, MMBench, and POPE, outperforming representative hybrid and discrete models under matched data, optimization, and parameter budgets.
Interpretation
It proposes a fully continuous unified multimodal paradigm: text blocks and images are mapped by modality-specific encoders into continuous embeddings that form hyperchunks preserving text token order and visual grid structure, and a single chunk-causal Flow Matching backbone learns one vector field over them. Prior unified models either quantize images into discrete visual tokens, creating a visual quantization bottleneck, or pair discrete language prediction with continuous image generation, requiring modality-dependent objectives and sampling procedures; this work lets both modalities share a continuous objective and sampling mechanism while retaining modality-specific representation statistics. The paper gives a full architecture description: frozen SigLIP 2 and T5-small encoders, a frozen image decoder and a separately trained text decoder, text chunks of length 8, an image as a single spatially structured chunk of 256 patch embeddings, plus the Flow Matching probability path and training objective.
It designs ordered hyperchunks and a chunk-causal backbone: predicting a chunk sees all preceding chunks while positions within the chunk are modeled jointly, multiple target chunks are predicted in parallel during training, and inference generates chunks sequentially with KV caching. Diverse multimodal tasks are expressed as task-defined ordered hyperchunk sequences, so the same backbone handles text-to-image, image-to-text, text-only, and image-only contexts without changing the model interface or objective. The paper provides the chunk-order factorization, the parallel training objective, sequence packing into at most 32,768 model positions, and the inference procedure that integrates the learned vector field chunk by chunk while caching keys and values of completed chunks.
It instantiates MF-1 and validates scaling, competitiveness, and transfer: under a common mixed-pretraining and evaluation protocol the 0.6B, 1.2B, and 1.6B variants show the Flow Matching objective and language perplexity decreasing while GenEval and CLIPScore increase; the 1.6B model scores 0.821 on GenEval and 83.44 on DPG-Bench, with an average of 75.3 across VQAv2, MMBench, and POPE. Trained from scratch on only 150B pretraining tokens, MF-1 surpasses from-scratch unified models such as Muddit and D-DiT and remains competitive with similarly sized models initialized from pretrained LLMs; under matched data, optimization, and trainable-parameter budgets it beats a Transfusion-style hybrid and a Chameleon-style fully discrete architecture on GenEval 0.7134, GQA 55.60, VQAv2 69.03, MMBench 46.74, and SEEDB 51.85. The paper reports a main results table for a 1.6B backbone with 150B pretraining tokens plus 5B joint finetuning tokens, and a controlled architecture comparison at 50B pretraining plus 5B finetuning where trainable parameter counts differ from MF-1 by at most 0.00156%.
Mixed pretraining yields transferable gains, and design analyses identify useful choices: at the 1.6B architecture and matched total training-token budgets, mixed pretraining raises SeedBench from 31.6 to 62.4 and MMBench from 36.0 to 67.2 relative to random initialization; semantic visual representations beat reconstruction-oriented VAE representations, modality-specific FFNs beat shared FFNs, and the same CFG mechanism improves both text-to-image and image-conditioned text generation. This indicates the benefit comes from mixed pretraining rather than increased token exposure, and presents the continuous visual representation space, the split between shared and modality-specific computation, and bidirectional CFG as reusable design choices. The paper includes a table contrasting random initialization and mixed pretraining on GenEval, DPG-Bench, SEEDB, MMB, and OK-VQA; a representation analysis comparing DINOv2 and SigLIP2 against FLUX.2 VAE latents, SD-VAE latents, and raw image patches; a table of four parameter-matched attention-projection and FFN configurations; and a 0.6B sweep where captioning peaks at guidance scale 3 and text-to-image at scale 5.
Perspective
The result targets unified multimodal pretraining and downstream finetuning over text and images: visual question answering and text-to-image generation are expressed as ordered hyperchunk sequences, with task data and optimization changing while the chunk interface, causal factorization, and continuous objective stay fixed. It lets researchers mix text-only, image-only, image-to-text, and text-to-image data in one backbone and reuse frozen modality codecs; for teams aiming to train unified models from scratch with fewer pretraining tokens, the 150B-token MF-1 setting plus released code and model provides a reproducible starting point. The authors list longer interleaved sequences, video, and other structured modalities as future directions.
Note that the main results table uses a single checkpoint with 150B pretraining tokens plus 5B joint finetuning, whereas the controlled architecture comparison and design analyses use a 50B pretraining setting, so cross-table comparisons should keep the differing budgets in mind. Language modeling is measured by GPT-2-large perplexity as an external fluency metric rather than the flow model's own likelihood, and captioning is measured by CIDEr and CLIPScore on the COCO Karpathy test split. The visual representation and CFG analyses use 0.6B models without downstream finetuning, so whether those conclusions hold at larger scale remains an open question. In addition, this reading is of the full text, but formulas and some tables appear as placeholders in the text, so specific numbers should be taken from the prose and tables.
