Public articles linked to the same research event.
arXiv Addressing why multimodal LLMs verify scientific claims substantially better from table evidence than from charts of the same underlying data, this work uses layer-wise linear probing and attention analysis on three open-weight VLMs and finds that chart information is encoded in intermediate representations but does not reach the prediction position, a disconnect absent for tables and taking two architecturally distinct forms across model families, reframing the table-chart gap as a failure of using encoded visual information at prediction time rather than a failure of encoding.
Addressing why multimodal LLMs verify scientific claims substantially better from table evidence than from charts of the same underlying data, this work uses layer-wise linear probing and attention analysis on three open-weight VLMs and finds that chart information is encoded in intermediate representations but does not reach the prediction position, a disconnect absent for tables and taking two architecturally distinct forms across model families, reframing the table-chart gap as a failure of using encoded visual information at prediction time rather than a failure of encoding.
Addressing why multimodal LLMs verify scientific claims substantially better from table evidence than from charts of the same underlying data, this work uses layer-wise linear probing and attention analysis on three open-weight VLMs and finds that chart information is encoded in intermediate representations but does not reach the prediction position, a disconnect absent for tables and taking two architecturally distinct forms across model families, reframing the table-chart gap as a failure of using encoded visual information at prediction time rather than a failure of encoding.
Addressing why multimodal LLMs verify scientific claims substantially better from table evidence than from charts of the same underlying data, this work uses layer-wise linear probing and attention analysis on three open-weight VLMs and finds that chart information is encoded in intermediate representations but does not reach the prediction position, a disconnect absent for tables and taking two architecturally distinct forms across model families, reframing the table-chart gap as a failure of using encoded visual information at prediction time rather than a failure of encoding.