Within an ε≤4/255 perturbation budget, targeted semantic substitution makes vision-language models name the target, confirm its presence, and deny the source, reaching 38% complete replacement on images and 35.9% on video
Related research and updatesSynopsis
Under a white-box threat model, the work aligns each stream of a source image with its counterpart in a target image in the victim vision-language model's post-merger token space and evaluates under a strict success criterion requiring the model to simultaneously name the target, confirm its presence, and deny the source, finding that target semantics appear at ε=2/255 and complete replacement reaches 38% at ε=4/255 on images and 35.9% at ε=1/255 on video, while also observing semantic fusion.
Figure 1: Qualitative image example on Qwen3-VL (dog → \rightarrow cat, ϵ \epsilon = 4/255). Adversarial images are omitted: at ϵ \epsilon = 4/255, the maximum per-pixel change is 1.6% of the full intensity range. All questions are posed independently to the VLM, each on the same adversarial image.
arXivInterpretation
It introduces a targeted semantic substitution attack on vision-language models that aligns corresponding streams of a source image and a target image in the post-merger token space, causing the model to replace source content with target semantics. Existing representation-alignment attacks achieve limited success at ε≤4/255, which suggested robustness in that range; this work shows targeted semantic substitution can succeed within the same range. Evaluated under a white-box threat model with a strict success criterion requiring the model to simultaneously name the target, confirm its presence, and deny the source; target semantics appear at ε=2/255 on images, complete replacement reaches 38% at ε=4/255 on images and 35.9% at ε=1/255 on video.
It demonstrates targeted semantic substitution on both image and video modalities and reports complete replacement rates at different perturbation budgets. Prior representation-alignment attacks had limited success at ε≤4/255; this work reports complete replacement on both images and video, extending the degree of semantic control achievable in that perturbation range. Complete replacement reaches 38% at ε=4/255 on images and 35.9% at ε=1/255 on video, both measured under the strict success criterion.
It observes a phenomenon of semantic fusion, in which the large language model rationalizes contradictory visual signals into a coherent narrative. This phenomenon reveals how the model behaves when faced with conflicting visual inputs, complementing the targeted semantic substitution results with an observation about internal integration. Reported as an experimental observation; the abstract provides no quantitative rate or statistical test.
Perspective
The work targets researchers and practitioners who need to evaluate the trustworthiness of vision-language models, and it applies to image and video settings under a white-box threat model with a perturbation budget of ε≤4/255. Its strict success criterion requires the model to simultaneously name the target, confirm its presence, and deny the source, so the reported complete replacement rates correspond to semantic control under this strong condition. The observation of semantic fusion offers an entry point for understanding how the model integrates conflicting visual signals and can inform subsequent work on internal integration mechanisms and robustness evaluation methods.
The abstract does not specify the identity or number of vision-language models used, the composition and scale of the image and video datasets, the statistical uncertainty of the complete replacement rates, or the quantitative extent of the semantic fusion phenomenon. How well the white-box results transfer to black-box or real deployment scenarios remains to be examined. The triggering conditions and stability of semantic fusion also warrant continued observation in follow-up work.
