PI3D places text-bearing objects in 3D scenes to make multiple MLLMs perform injected tasks, and shows existing defenses fail to reliably stop it
Related research and updatesSynopsis
The work introduces PI3D, a prompt injection attack against multimodal large language models (MLLMs) in 3D environments: instead of digitally editing images, the attacker places a text-bearing 3D object in the physical environment, and the paper formulates and solves the problem of finding an effective pose (position and orientation) that induces the MLLM to perform the injected task while keeping the placement physically plausible; experiments show PI3D is effective against multiple MLLMs under diverse camera trajectories, and that a range of evaluated defenses is not sufficient to reliably defend against it.
Figure 1: Prompt injection in 3D environments: a whiteboard with injected text causes the MLLM to mistakenly describe the outdoor suburban neighborhood as a “library.”
arXivInterpretation
It proposes PI3D, extending prompt injection from the text domain and digitally edited 2D images to 3D physical environments, where the attack carrier is a text-bearing physical object rather than modified image pixels. Prior work studied prompt injection in the text domain and through digitally edited 2D images, while limited attention was paid to how such attacks function in 3D environments; PI3D addresses this gap by targeting camera-captured views of the physical world. The paper presents PI3D as a problem formulation and attack construction, stating that it is realized through text-bearing object placement rather than digital image edits.
It formulates the attack as a pose-solving problem: finding an effective position and orientation for a 3D object with injected text so that the model performs the injected task while the object placement remains physically plausible. By making attack effectiveness and physical plausibility joint objectives, the injection no longer relies on arbitrarily editable image content but is constrained by how objects can actually be placed in a real scene. The abstract states that the authors 'formulate and solve' this pose identification problem, with both the induced-task goal and the physical-plausibility requirement.
Experiments show PI3D is an effective attack against multiple MLLMs under diverse camera trajectories. Effectiveness is not limited to a single model or a single viewpoint, indicating the attack holds across different models and different observation paths. The abstract reports that 'PI3D is an effective attack against multiple MLLMs under diverse camera trajectories', a multi-model, multi-trajectory experimental conclusion.
Evaluating a range of defenses shows they are not sufficient to reliably defend against PI3D. The evaluation extends from attack success to the defense side, indicating that current defenses struggle to provide reliable protection under this threat model. The abstract states 'we further evaluate a range of defenses and show that they are not sufficient to reliably defend against PI3D', a defense-evaluation-level conclusion.
Perspective
The work targets multimodal large language models that reason and act on camera-captured views in 3D environments, with application settings including robotics and situated conversational agents. Its attack setting assumes the attacker can place a text-bearing 3D object in the physical environment, treats pose (position and orientation) as the controllable variable, and requires the placement to remain physically plausible. The conclusions cover attack effectiveness across multiple models and camera trajectories and the insufficiency of the evaluated defense set; no conclusions are offered for models, environment types, or defense mechanisms outside that evaluation.
The visible text is at the abstract level and does not give attack success rates, the specific pose-search algorithm, the list of models used, the number of camera trajectories, the defense types, or evaluation metrics, so the magnitude of effects and differences across conditions cannot be judged. The criteria for physical plausibility, the attacker's prior knowledge requirements about the environment, and whether defenses remain insufficient under stronger settings are open questions that require the full paper. The abstract also does not describe behavior in dynamic environments or human-occupied scenes.
