Skip to main content
Back to timeline
arXivSource publication:

Multimodal LLMs encode chart information but fail to route it to the prediction position, as layer-wise probing and attention analysis explain the table-chart verification gap

Related research and updates

Synopsis

Addressing why multimodal LLMs verify scientific claims substantially better from table evidence than from charts of the same underlying data, this work uses layer-wise linear probing and attention analysis on three open-weight VLMs and finds that chart information is encoded in intermediate representations but does not reach the prediction position, a disconnect absent for tables and taking two architecturally distinct forms across model families, reframing the table-chart gap as a failure of using encoded visual information at prediction time rather than a failure of encoding.

Source-provided article image: Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification
Figure 1 ·

Figure 1: Overview of our approach. We apply linear probing to compare how chart and table evidence is encoded and passes to the model response.

arXiv

Interpretation

The work answers whether models fail to extract information from charts or extract it but fail to use it, and finds consistent evidence for the latter: chart information is encoded in the models' intermediate representations but does not reach the prediction position. Prior work had shown only that models perform substantially better with table evidence than with charts of the same underlying data; this study moves from that performance gap to a mechanistic account separating encoding from prediction. Evidence comes from layer-wise linear probing and attention analysis on three open-weight VLMs, with tables and charts representing the same underlying data, and the finding holds across all conditions tested.

Tables do not show this disconnect between encoding and prediction, indicating the phenomenon is tied to how evidence is presented rather than a general inability to use intermediate representations. By contrasting tables and charts over the same underlying data, the gap is localized to presentation form rather than data content. The comparison is conducted on the same underlying data and task, with the disconnect absent for tables and consistently present for charts.

Attention analysis further reveals that this disconnect takes two architecturally distinct forms across model families. It refines a single performance gap into distinguishable mechanistic types across model families, suggesting the issue is not specific to one implementation. The conclusion rests on attention analysis of three open-weight VLMs, with the disconnect differing in form across the model families tested.

Perspective

The result applies to scientific claim verification where evidence is a table or a chart and the task is to judge whether a claim is supported, especially in multimodal LLM assistance for peer review. It speaks to researchers studying how multimodal models use evidence and to practitioners designing verification pipelines that require models to read chart evidence. In that setting, the findings suggest improvement efforts can target how encoded visual information is called upon at prediction time, rather than only making chart extraction better.

Readers may still watch which model families correspond to each of the two architecturally distinct forms of the disconnect and how stable that distinction is across more models; under what conditions the disconnect weakens or disappears; and how much verification performance recovers once prediction-time use of visual information is improved. The loaded text is at the abstract level and does not include the specific probing setup, attention metrics, or per-model results, so those details remain to be confirmed in the paper's figures and experimental sections.

Sources