Skip to main content
Back to timeline
arXivSource publication:

OmniTaskonomy maps 19 generation tasks against 25 understanding capabilities: I2I-then-I2T training improves understanding, and gradient alignment tracks the gains

Synopsis

Using controlled paired image-to-image (I2I) generation and image-to-text (I2T) understanding tasks on the unified multimodal model BAGEL, the authors find that a recipe of I2I training that updates shared parameters followed by I2T finetuning improves understanding with gains that grow as I2I data increases, build OmniTaskonomy, a taxonomy spanning 19 I2I tasks and 25 understanding capabilities, produce a selective task-dependent generation-to-understanding transfer map, and find that stronger gradient alignment between generation and understanding is associated with larger downstream transfer gains.

AI-generated editorial illustration: OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Interpretation

On paired tasks, I2I training improves the corresponding I2T understanding performance, with gains growing as I2I data increases; directly mixing the two objectives or freezing shared parameters yields weaker or less stable gains. Prior work largely showed understanding helping generation, while the reverse was unclear and sometimes argued to be limited; this work turns that asymmetry into a measurable training-recipe question using paired tasks that share the same input and underlying visual problem and differ only in output modality. Six training recipes are compared on BAGEL-7B-MoT for Jigsaw and Zoom-In, with the I2T budget fixed at 1k examples and the amount of I2I data varied; each result is the mean over three random seeds with standard error reported, and I2I then I2T is the default recipe.

I2I supervision can partially substitute for direct I2T supervision, with larger benefits when understanding data is limited. This quantifies the value of generation data when understanding data is scarce, rather than only giving a directional conclusion. With the I2I budget fixed at 100k examples and compared against I2T-only at 1k, 3k, 10k, and 30k budgets: on Zoom-In, 100k I2I plus only 1k I2T reaches performance comparable to 10k I2T alone, and on Jigsaw, 100k I2I plus only 3k I2T reaches performance comparable to 10k I2T alone.

OmniTaskonomy places 19 I2I tasks and 25 understanding capabilities in one capability-based taxonomy, and the transfer map is highly non-uniform: metric 3D relation and counting benefit from the broadest set of sources (12 I2I tasks each), 2D ordering follows with eleven, while OCR, text recognition, and appearance understanding show no significant improvement from any tested source. Existing benchmarks organize visual tasks by task formulation or output modality, leaving generation and understanding separated; this taxonomy aligns both under shared visual capabilities (Recognition, Reconstruction, Reorganization), enabling cross-modal transfer to be measured at the level of individual capabilities. The taxonomy is derived bottom-up from samples of seven vision-language benchmarks, retained 9,444 examples by majority vote of three VLM judges (93.30% unanimous, 6.70% two-of-three), and was validated by four human reviewers on 250 samples, with 97.4% of definite human judgments matching the VLM label.

Transfer appears both between closely matched capabilities and in less direct pairings, and gradient alignment is positively associated with transfer gains. This offers a testable optimization-level signal for which generation tasks to select, beyond an empirical task-pairing table. In controlled pairs, I2I and I2T gradients align most strongly in the understanding branch's pre-attention RMSNorm parameters, particularly in earlier Transformer layers; across tasks, average alignment and average transfer gain are positively correlated across seven capabilities, and alignment and gain are positively correlated across the 133 source-target pairs, with transfer significance assessed by paired permutation tests.

Perspective

The results target unified multimodal models that learn generation and understanding within the same model, especially the Mixture-of-Transformers setting where generation and understanding branches share attention; they indicate that when understanding data is limited, one can first use generation training to establish a useful initialization, then select generation tasks for the target capabilities, and use gradient alignment as a signal for identifying promising task pairings. The transfer map covers 19 I2I tasks and 25 understanding capabilities, evaluated on 9,444 samples from seven benchmarks, so its intended setting is visual understanding evaluation within these capabilities and tasks.

The transfer map is highly non-uniform, and OCR, text recognition, and appearance understanding show no significant improvement from any tested source, so which capabilities are intrinsically hard to benefit from generation supervision remains open. Gradient alignment and transfer gain are positively associated, but the text does not establish a causal mechanism, so whether alignment can serve as an ex-ante screening tool still needs prospective validation. The controlled paired experiments cover only Jigsaw and Zoom-In, and the cross-task analysis uses a fixed budget of roughly 50k I2I plus 50k LLaVA-Instruct examples, leaving behavior at larger data scales and in other architectures to be examined. The taxonomy relies on VLM annotation and human review; although agreement is high, the boundaries between capabilities still involve definitional choices.

Sources