Public articles linked to the same research event.
arXiv MLCommons Jailbreak Benchmark v1.0 introduces an end-to-end methodology that evaluates eight open-weight large language models with paired baseline and adversarial conditions using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy, judging responses with the AILuminate Assessment Standard v1.4 and measuring robustness through the Resilience Gap; the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57%, with accessible systems showing a larger mean gap and attack effectiveness varying substantially across attack categories and hazards.
MLCommons Jailbreak Benchmark v1.0 introduces an end-to-end methodology that evaluates eight open-weight large language models with paired baseline and adversarial conditions using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy, judging responses with the AILuminate Assessment Standard v1.4 and measuring robustness through the Resilience Gap; the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57%, with accessible systems showing a larger mean gap and attack effectiveness varying substantially across attack categories and hazards.
MLCommons Jailbreak Benchmark v1.0 introduces an end-to-end methodology that evaluates eight open-weight large language models with paired baseline and adversarial conditions using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy, judging responses with the AILuminate Assessment Standard v1.4 and measuring robustness through the Resilience Gap; the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57%, with accessible systems showing a larger mean gap and attack effectiveness varying substantially across attack categories and hazards.
MLCommons Jailbreak Benchmark v1.0 introduces an end-to-end methodology that evaluates eight open-weight large language models with paired baseline and adversarial conditions using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy, judging responses with the AILuminate Assessment Standard v1.4 and measuring robustness through the Resilience Gap; the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57%, with accessible systems showing a larger mean gap and attack effectiveness varying substantially across attack categories and hazards.