Skip to main content
Back to timeline
arXivSource publication:

MIST stress test: adding an image shifts about a fifth of VLM judge labels, regardless of what the image shows

Synopsis

The authors built MIST (the Misleading-Image Stress Test), 200 English sentences each containing a phrase readable either figuratively or literally and shown with an aligned image, a misleading image, or no image; across thirteen VLM judges an aligned image changed 20.5% of labels and a misleading one 19.4%, versus 11.6% when only the ignore-the-image instruction was deleted with the image left in place, and only 37% of the labels that differ between the two images moved toward the sense shown, while agreement with the human annotators was unchanged whether the image was absent, aligned, or misleading.

AI-generated editorial illustration: It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them

Interpretation

MIST separates whether an image is used from whether its content influences the judgment: aligned and misleading images move nearly the same share of labels (20.5% versus 19.4% across thirteen judges; 15.9% versus 14.7% for the seven judges reported in the body, never more than five points apart for any judge). Earlier multimodal benchmarks such as IRFL and AdMIRe make the image the object of the decision, so changing it legitimately changes the answer; MIST requires the label to come from the sentence alone, making the image a label-preserving perturbation by construction, so any label change is an error. 200 English items drawn from 200 distinct compounds, fifty per condition; thirteen judges, four prompts, explicit and silent image instructions, 52,000 labels in total, greedily decoded and constrained to the four labels.

The changes do not follow the image: of the 17.1% of cells (1,776 of 10,400) where the two images disagree, only 37% move toward the sense the misleading image depicts and 63% move away; every judge that passes the alt-test falls below the 50% chance level, the highest at 39%. This contradicts the expectation that each image pulls a judge toward the sense it depicts, locating the effect in the presence of an image rather than in what it shows. The imbalance holds for each sentence type separately (683 figurative cells, 38% following the image; 1,093 literal cells, 35%), and 71% of the movement stays on the same side of the figurative–literal divide.

The image does not reduce agreement with the human majority: 54.4% with no image, 53.7% aligned, 54.7% misleading, and only 5 of 13 judges lose accuracy under an image. Aggregate agreement cannot see instance-level instability: each change is an error by construction, but they cancel in the aggregate, so a substitutability verdict computed from overall agreement is blind to them. Agreement is exact match with the human majority label of the trio that saw the item; agreement is lower on misleading items than aligned ones, but that gap is largest with no image at all, pointing to the difficulty of the two disjoint expression sets.

The instability sits in the models a practitioner would deploy: the seven judges that pass the alt-test move 15.9% of labels under an image against 25.8% for the six that never pass, yet each group exceeds the control that deletes only the ignore-the-image instruction (8.8% for the seven, 11.6% across all thirteen). It reframes a substitutability verdict as describing a judge and a configuration at once, and recommends that a deployment report state the prompt used and any irrelevant context attached. The control holds the image fixed and deletes only the paragraph instructing the judge to ignore it, giving an upper bound on how much a prompt edit alone can move a judge; every judge individually exceeds it.

Perspective

The work addresses evaluation and data-production pipelines that use VLM judges in place of human annotators: it releases MIST, an invariance test, together with all human and VLM labels, so practitioners can check whether a verdict holds under their own prompt and context configuration and can state the prompt used and any irrelevant context attached in a deployment report. The results apply to English, to the potentially idiomatic expression task, and to judges run with reasoning modes disabled.

Both images depict a reading of the target phrase, so the contrast is between two relevant images rather than a relevant against an irrelevant one; the labels are ordered FF to LL, but the FF-versus-WF and FL-versus-LL boundaries are different judgments, so a one-step move is not the same quantity everywhere; because each human annotator saw one condition per expression, there is no human change rate, and the prompt-edit control supplies only a within-judge baseline; the human aligned-versus-misleading contrast is not statistically resolvable, and the authors report it as unresolved; whether judges also move when the sentence genuinely changes reading is the complementary measurement MIST does not make; and whether the effect is specific to the alt-test and figurative annotation is the next question.

Sources