Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

PI3D places text-bearing objects in 3D scenes to make multiple MLLMs perform injected tasks, and shows existing defenses fail to reliably stop it

The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.