UniData turns a simple requirement into multi-round, nine-modality instructions, beating Self-Instruct and VIGC on reasonableness, clarity, detail and relevance
Synopsis
The authors introduce UniData, a universal multimodal instruction generation pipeline that first expands user keywords into diverse events, then uses a LLaMA-3-backed multimodal instruction generator with image and music tokenizers to produce instructions round by round, correcting repetition and irrelevance at inference via inter-round correlation; they also build UniDataset with nine modalities, 20,000 entries and 17.5 rounds on average, and report that its data quality exceeds Self-Instruct and VIGC on all four GPT-guided metrics while fine-tuning on the generated data improves other multimodal models' understanding and generation.
Figure 1: Usage of UniData. (I) Users only input one simple requirement, combined with random emotions and occasions to create diverse events. (II) Each diverse event will generate one multimodal instruction. (III) The generated multimodal instructions include various modalities.
arXivInterpretation
UniData generates multi-round, multimodal instructions from one simple requirement, covering nine modalities on both input and output sides with 17.5 rounds per dialogue on average, whereas Self-Instruct and VIGC support only language or language-plus-image with fewer rounds. Self-Instruct was confined to the language modality, VIGC added visual understanding, and Multimodal Self-Instruct produced only abstract images; UniData widens modality coverage to nine and supports long multi-round generation. Tables 1 and 10 report the comparison: UniData averages 1.92 images, 0.74 music clips, 7.37 other modalities and 17.5 rounds, while the baselines show empty image and music columns and 3 and 1 rounds.
On the four GPT-guided quality metrics, UniData surpasses both baselines on every dimension, with the largest margins on clarity and relevance. The baselines had not previously been compared under these four metrics on the same test data; UniData supplies that side-by-side comparison. Table 1 gives UniData 0.661 reasonableness, 0.527 clarity, 0.667 detail and 0.682 relevance, versus Self-Instruct at 0.624/0.444/0.594/0.520 and VIGC at 0.460/0.359/0.534/0.569; the 15-evaluator forced-choice user study in Appendix C.3 shows UniData preferred on every dimension.
Fine-tuning other multimodal models on UniData-generated data improves visual understanding, image understanding and generation, music, code, math and text-based benchmarks. This tests the transferable value of the generated data for downstream models rather than only the quality scores of the data itself. Table 7 shows MMMU 35.8 to 36.4, ChartQA 17.4 to 17.9, TextVQA 58.9 to 60.1, MBPP 20.8 to 21.3, MathVista 59.4 to 60.1, MathVision 12.8 to 13.0 and MMLU-Pro 20.3 to 21.9; image understanding in the interleaved setting rises from 19.42% to 35.25%; music shows lower FAD and KL with higher CLAP, and the authors note the music gains are small.
Ablations show event expansion and instruction-flow correction each contribute to quality, with all four metrics dropping when either is removed. The two pipeline design choices are tied directly to final data quality in a controlled comparison. In Table 7, w/o Expansion scores 0.557/0.518/0.607/0.550 and w/o Correction 0.648/0.489/0.610/0.654, against 0.661/0.527/0.667/0.682 for the full method; the authors also report that without expansion the number of generatable instructions drops substantially.
Perspective
The work targets MLLM training and data-production settings that need large-scale multimodal instruction data: a user supplies a few keywords and obtains a multi-round, nine-modality instruction set, and can steer it toward a chosen modality such as image or music (Table 14 shows image-related input raising image count to 14.90 and music-related input raising music count to 12.30). The authors position UniData as a text-centered orchestration pipeline rather than a native any-to-any model, so per-modality quality is bounded by the external tools. The pipeline is designed to absorb new modalities: Stage 3 emits modality-agnostic slots and Stage 4 can plug in new APIs through a thin adapter. The authors list follow-up directions including replacing the backbone with larger, newer models, adding video, 3D, motion and tabular generation tools, and letting users supply a downstream specification for purpose-built data.
Quality scoring relies on GPT-4o as judge, and how well its preferences align with human judgment remains an open question despite the 15-evaluator user study. Downstream gains are limited on most benchmarks, for example MathVision 12.8 to 13.0 and MMMU 35.8 to 36.4, and whether these hold for larger models or longer training is not settled in the text. Instruction-flow correction currently operates only at the textual-prompt level, and the authors note a tighter loop conditioning the corrector on rendered media is left to future work, so whether repetition or drift persists in the final media is worth watching. The authors also note the nine-modality coverage still omits video, 3D, motion, tables, biological data and spoken input, and that the pipeline inherits biases from its LLM backbones, search APIs and modality generators, with systematic bias audits remaining future work.
