Public articles linked to the same research event.
arXiv The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.
The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.
The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.
The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.