Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

Study reveals two fusion pathways in multimodal LLMs: concatenation models go text-first then vision, native models show early visual-textual co-adaptation

The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.