Only content delivered through the publication boundary on this date is included.
arXiv The work proposes JMLLM, a jailbreak framework that combines four concealment strategies—alternating translation, word encryption, feature collapse, and harmful injection—across text, visual, and speech modalities, releases the TriJail dataset with 1,250 textual and speech adversarial prompts plus 150 harmful images, reports leading attack success rates on 16 popular LLMs over AdvBench and TriJail at 24.65 seconds for a single query and 6 queries for the multi-round mode, and proposes a Harmful Separator defense that reduces attack success rates.
The work proposes JMLLM, a jailbreak framework that combines four concealment strategies—alternating translation, word encryption, feature collapse, and harmful injection—across text, visual, and speech modalities, releases the TriJail dataset with 1,250 textual and speech adversarial prompts plus 150 harmful images, reports leading attack success rates on 16 popular LLMs over AdvBench and TriJail at 24.65 seconds for a single query and 6 queries for the multi-round mode, and proposes a Harmful Separator defense that reduces attack success rates.
The work proposes JMLLM, a jailbreak framework that combines four concealment strategies—alternating translation, word encryption, feature collapse, and harmful injection—across text, visual, and speech modalities, releases the TriJail dataset with 1,250 textual and speech adversarial prompts plus 150 harmful images, reports leading attack success rates on 16 popular LLMs over AdvBench and TriJail at 24.65 seconds for a single query and 6 queries for the multi-round mode, and proposes a Harmful Separator defense that reduces attack success rates.
The work proposes JMLLM, a jailbreak framework that combines four concealment strategies—alternating translation, word encryption, feature collapse, and harmful injection—across text, visual, and speech modalities, releases the TriJail dataset with 1,250 textual and speech adversarial prompts plus 150 harmful images, reports leading attack success rates on 16 popular LLMs over AdvBench and TriJail at 24.65 seconds for a single query and 6 queries for the multi-round mode, and proposes a Harmful Separator defense that reduces attack success rates.