Public articles linked to the same research event.
arXiv Using a controlled two-stage PGD attack on 200 held-out COCO images with Qwen2.5-VL-7B-Instruct, the authors show that a targeted perturbation can drive teacher-forced training loss for a fixed target caption to near zero while free generation still produces the original correct description, localize this train/inference gap to a single autoregressive step (target-token rank fixed at exactly 3,488 of 152,064 vocabulary entries with zero variance), and trace all 28 decoder layers to show the visual encoder corrupts every image by a comparable margin while the language decoder amplifies the corrupted signal for susceptible images and actively suppresses it for resistant ones (p<0.001, rank-biserial r=0.579).
Using a controlled two-stage PGD attack on 200 held-out COCO images with Qwen2.5-VL-7B-Instruct, the authors show that a targeted perturbation can drive teacher-forced training loss for a fixed target caption to near zero while free generation still produces the original correct description, localize this train/inference gap to a single autoregressive step (target-token rank fixed at exactly 3,488 of 152,064 vocabulary entries with zero variance), and trace all 28 decoder layers to show the visual encoder corrupts every image by a comparable margin while the language decoder amplifies the corrupted signal for susceptible images and actively suppresses it for resistant ones (p<0.001, rank-biserial r=0.579).
Using a controlled two-stage PGD attack on 200 held-out COCO images with Qwen2.5-VL-7B-Instruct, the authors show that a targeted perturbation can drive teacher-forced training loss for a fixed target caption to near zero while free generation still produces the original correct description, localize this train/inference gap to a single autoregressive step (target-token rank fixed at exactly 3,488 of 152,064 vocabulary entries with zero variance), and trace all 28 decoder layers to show the visual encoder corrupts every image by a comparable margin while the language decoder amplifies the corrupted signal for susceptible images and actively suppresses it for resistant ones (p<0.001, rank-biserial r=0.579).
Using a controlled two-stage PGD attack on 200 held-out COCO images with Qwen2.5-VL-7B-Instruct, the authors show that a targeted perturbation can drive teacher-forced training loss for a fixed target caption to near zero while free generation still produces the original correct description, localize this train/inference gap to a single autoregressive step (target-token rank fixed at exactly 3,488 of 152,064 vocabulary entries with zero variance), and trace all 28 decoder layers to show the visual encoder corrupts every image by a comparable margin while the language decoder amplifies the corrupted signal for susceptible images and actively suppresses it for resistant ones (p<0.001, rank-biserial r=0.579).