Skip to main content
Back to timeline
arXivSource publication:

Study reveals two fusion pathways in multimodal LLMs: concatenation models go text-first then vision, native models show early visual-textual co-adaptation

Related research and updates

Synopsis

The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.

Source-provided article image: Architecture-Dependent Fusion Pathways in MLLMs
Figure 1 ·

Figure 1: Cross-modal CKA similarity across layers. Family average represents the average values of different types of models. Shaded regions mark the early, transition, and late layer ranges. The model classification is described in EXPERIMENTAL SETUP 3.1 . Model-name colors distinguish concatenation from native architectures.

arXiv

Interpretation

The work identifies two distinct fusion pathways: concatenation-architecture models follow a text-first, vision-later pathway, whereas native multimodal models exhibit earlier visual-textual co-adaptation and feature-space reorganization. Prior understanding of how visual and textual information is fused across layers inside MLLMs was insufficient; this work characterizes the two architecture families mechanistically in terms of fusion timing and feature-space organization. Based on progressive analyses of representative models from both architectures: alignment decoupling identifies which modality changes, attention routing and entropy characterize cross-modal information distribution, and intrinsic dimensionality examines how fusion reshapes feature spaces; causal intervention experiments separately validate the resulting interpretation.

The three progressive analyses are connected, respectively addressing which modality changes, how cross-modal information is distributed, and how fusion reshapes feature spaces. It links modality alignment, attention distribution, and feature-space geometry into a single analysis chain rather than examining a single metric in isolation. The abstract explicitly lists the three analyses and their targets, and states that causal intervention experiments are performed separately to validate the resulting interpretation.

As a supplementary analysis, the study uses visual CKA to examine the Platonic Representation Hypothesis. It brings a representation-similarity measure into the mechanistic discussion of multimodal fusion as a complementary perspective to the main analyses. The abstract positions this as a supplementary analysis and does not report specific values or conclusions for this part.

Perspective

The work targets researchers and practitioners studying internal fusion mechanisms in multimodal large language models, and applies to architecture-aware diagnostics of concatenation and native multimodal architectures. Its analysis framework takes representative models from the two architectures as its object, and the conclusions apply within the scope of the models examined; causal intervention experiments validate the interpretation derived from the analyses, and visual CKA serves as a supplementary analysis of the Platonic Representation Hypothesis. Readers can use it to anticipate different fusion timing and feature-space organization by architecture type when designing or diagnosing multimodal models.

From the abstract alone, the specific list of models examined, datasets, sample sizes, and effect sizes are unavailable, as are the concrete design of the causal intervention experiments and the quantitative results of the visual CKA analysis. How the three analyses connect, the specific measure of intrinsic dimensionality, and the outcome of the Platonic Representation Hypothesis test all require the full text. In addition, the stability of the two fusion pathways across model scales, training objectives, and tasks remains an open question worth watching.

Sources