Qwen2.5-VL-7B shows a train/inference gap on 200 COCO images: target-token rank is fixed at 3,488 per step, and the language decoder, not the visual encoder, decides whether the attack lands
Related research and updatesSynopsis
Using a controlled two-stage PGD attack on 200 held-out COCO images with Qwen2.5-VL-7B-Instruct, the authors show that a targeted perturbation can drive teacher-forced training loss for a fixed target caption to near zero while free generation still produces the original correct description, localize this train/inference gap to a single autoregressive step (target-token rank fixed at exactly 3,488 of 152,064 vocabulary entries with zero variance), and trace all 28 decoder layers to show the visual encoder corrupts every image by a comparable margin while the language decoder amplifies the corrupted signal for susceptible images and actively suppresses it for resistant ones (p<0.001, rank-biserial r=0.579).
Figure 1 : Per-token probability analysis across 200 images and all three outcome groups. The rank of the target token _coconut , conditioned on the correct preceding token A already appearing in context, is a flat line at exactly 3,488 with zero variance, for every image, every group, clean and adversarial alike.
arXivInterpretation
The paper names and characterizes a train/inference gap: a targeted adversarial perturbation can drive a VLM's teacher-forced training loss for a fixed target caption to near zero, yet the same model generating freely still produces the original correct description with no trace of the target. Prior adversarial-robustness discussion largely focused on whether an attack changes the output; this work explicitly separates 'loss driven down' from 'generation unchanged' as two dissociable phenomena and sets a mechanistic explanation target. Verified on Qwen2.5-VL-7B-Instruct with a controlled two-stage PGD attack on 200 held-out COCO images, a single-model, single-dataset-scale mechanistic study.
Image-level pixel statistics have essentially no predictive power over which images get corrupted, including a correctly re-implemented texture-based attackability measure from the CNN robustness literature. Transfers attackability predictors common in the CNN era to the VLM setting and tests their failure, showing susceptibility is not explained by surface statistics of the input image. Best predictor r=-0.050, p=0.484; ridge regression R^2=0.069, both indicating near-chance predictive power.
The gap is localized to a single autoregressive step: conditioned on the correct first token already being generated, the rank of the target token is exactly 3,488 out of 152,064 vocabulary entries for every image and every condition, with zero variance. Uses the logit lens to collapse a diffuse phenomenon into one deterministic, reproducible step, giving subsequent mechanistic analysis and defense targeting a clear coordinate. The rank value is identical across all images and conditions with zero variance, a strong regularity observation, though based on a single model and fixed attack setup.
Tracking all 28 LLM decoder layers shows the visual encoder corrupts every image's representation by a comparable margin, while the language model decoder differentially arbitrates: amplifying the corrupted signal for susceptible images and actively suppressing it past its clean-image baseline for resistant ones. Shifts the attribution of robustness from the visual encoder to the language decoder's prior, indicating defenses and faithfulness evaluations should target the decoder rather than only the visual front end. Between-group difference p<0.001, rank-biserial r=0.579; a linear probe on the merger hidden state separates the two outcomes with AUC=0.858, though the authors flag a circularity concern in this estimate.
Perspective
This work targets adversarial robustness and faithfulness evaluation in autoregressive vision-language models, with experiments set on Qwen2.5-VL-7B-Instruct and a controlled two-stage PGD attack on 200 held-out COCO images, suited to researchers and defense designers who need to understand why training loss and generation behavior decouple. Its conclusion directly points to defenses and evaluations targeting the language decoder prior rather than only the visual encoder, offering a reproducible starting point for testing this mechanism on other VLMs, other attacks, and larger image scales.
Readers should still watch: the mechanistic account rests on a single model and 200 COCO images, so whether the same fixed rank and decoder-arbitration pattern holds across other VLM architectures, other attack methods, and larger data scales remains to be tested; the circularity concern in the merger-hidden-state linear probe with AUC=0.858 is flagged by the authors themselves, and independent confirmation of its discriminative power requires follow-up work; moreover, the cause behind the zero-variance fixed target-token rank of 3,488, and whether it is tied to a specific vocabulary, prompt format, or attack hyperparameters, remain open questions.
