Skip to main content
Back to timeline
arXivSource publication:

Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models

Synopsis

The work proposes JMLLM, a jailbreak framework that combines four concealment strategies—alternating translation, word encryption, feature collapse, and harmful injection—across text, visual, and speech modalities, releases the TriJail dataset with 1,250 textual and speech adversarial prompts plus 150 harmful images, reports leading attack success rates on 16 popular LLMs over AdvBench and TriJail at 24.65 seconds for a single query and 6 queries for the multi-round mode, and proposes a Harmful Separator defense that reduces attack success rates.

Source-provided article image: Divide and Conquer: A Hybrid Strategy Defeats Multimodal Large Language Models
Fig. 1 ·

Fig. 1: A typical application scenario illustrating the potential harms of LLM jailbreaks.

arXiv

Interpretation

It proposes JMLLM, described as the first hybrid-strategy jailbreak framework covering text, visual, and speech modalities, using alternating translation, word encryption, feature collapse, and harmful injection to bypass per-modality defenses. Prior methods mostly target a single modality or a text-plus-visual pair; this work integrates three modalities in one framework and separates single-query and multi-query attack modes. The paper gives formula-level descriptions and pseudocode for all four strategies (Algorithms 1 and 2) and reports ablations on TriJail and AdvBench showing ASR drops when any module is removed.

It builds the TriJail dataset with 1,250 text prompts, 1,250 speech prompts, and 150 harmful images across six scenarios: Hate Speech and Discrimination, Misinformation and Disinformation, Violence, Threats, and Bullying, Pornographic Exploitative Content, Privacy Infringement, and Self-Harm. The paper states existing jailbreak datasets are single- or bi-modal and concentrate on limited domains such as bombs, drugs, and violence; TriJail adds the speech modality and a broader scenario taxonomy. Table I reports per-scenario counts of texts, images, speech, words, and tokens; construction involved manual extraction and rewriting from forums, manual prompt design, then TTS-1 for speech and DALL-E-3 for images.

On AdvBench, JMLLM-Single exceeds GCG, AutoDAN, PAIR, and ReNeLLM in both GPT-ASR and KW-ASR on GPT-3.5-turbo, GPT-4, Claude-1, Claude-2, and Llama2-7B, while taking 24.65 seconds per attack versus 132.03 seconds for ReNeLLM. Relative to the strong ReNeLLM baseline, the paper reports roughly 5.36 times faster execution with higher ASR; the multi-round mode reaches further gains with only 6 queries. Table V lists GPT-ASR, KW-ASR, per-attack time, and query counts for each baseline; Table XIII further compares JMLLM and ReNeLLM across seven AdvBench scenarios.

It proposes the Harmful Separator defense, which splits a jailbreak prompt into a harmless instruction and a potentially harmful example and analyzes the example separately, reducing JMLLM ASR by 0.622, 0.385, and 0.547 on GPT-3.5-turbo, Llama-3.1-405B, and GPT-4o respectively. The paper reports that a Useful and Safe instruction method barely reduces attack success, whereas Harmful Separator substantially lowers it without eliminating the risk. Table XIV reports baseline JMLLM ASR and the ASR change under two defenses across three models; the defense is inspired by structured-query frontend work.

Perspective

The results target multimodal LLM safety evaluation and red-teaming settings with text, visual, and speech inputs; TriJail covers six prohibited scenarios and supports comparisons of attack success rates across models and defenses. The paper notes that current multimodal models are expanding to video, haptic, and other modalities, so jailbreak research needs to extend to more complex multimodal environments and embodied agents; the conclusions therefore apply mainly to the modalities and models tested here.

The paper reports that TOX-ASR is markedly lower than the other three metrics and that KW-ASR may be inflated because it only checks for keywords, so success rates differ by evaluation criterion; the speech experiments randomly draw 20 adversarial samples per scenario, and the sample size and statistical variability of the visual and speech ablations are not expanded in the main text; the paper also notes the black-box problem of closed-source models, the risk that attacks become obsolete after model updates, and the compliance and ethical costs of collecting jailbreak data, which are directions for continued observation.

Sources