Public articles linked to the same research event.
arXiv The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.
The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.
The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.
The study runs three connected analyses (alignment decoupling, attention routing and entropy, intrinsic dimensionality) plus causal intervention on representative concatenation and native multimodal LLMs, finding that concatenation models follow a text-first, vision-later fusion pathway while native models show earlier visual-textual co-adaptation and feature-space reorganization, offering a mechanistic view of multimodal fusion and architecture-aware diagnostics.