Skip to main content
Back to timeline
arXivSource publication:

MLCommons releases Jailbreak Benchmark v1.0: unsafe-response rate across eight open-weight models rises from 11.08% to 18.65% under jailbreak conditions, with an average Resilience Gap of 7.57%

Related research and updates

Synopsis

MLCommons Jailbreak Benchmark v1.0 introduces an end-to-end methodology that evaluates eight open-weight large language models with paired baseline and adversarial conditions using 264 seed prompts spanning eleven hazard categories and representative attacks drawn from the MLCommons Jailbreak Taxonomy, judging responses with the AILuminate Assessment Standard v1.4 and measuring robustness through the Resilience Gap; the unsafe-response rate rose from 11.08% at baseline to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57%, with accessible systems showing a larger mean gap and attack effectiveness varying substantially across attack categories and hazards.

Source-provided article image: MLCommons Jailbreak Benchmark v1.0
Figure 1 ·

Figure 1: Resilience gap by SUT: safe-response rates before and after the application of jailbreak techniques.

arXiv

Interpretation

The benchmark provides an end-to-end, reproducible jailbreak evaluation pipeline that combines criteria-driven system and attack selection, paired baseline and adversarial evaluation, human annotation, automated evaluator calibration, scoring, grading, and risk-calibrated disclosure. Relative to fragmented prior jailbreak evaluation practice, it places the assessment standard, attack taxonomy, and disclosure rules within a single methodological framework, explicitly using the AILuminate Assessment Standard v1.4 as the basis for judging responses. The methodological description comes from the abstract and covers each stage of the evaluation pipeline; implementation details, code, and data are not expanded in the abstract.

Across eight open-weight systems, 264 seed prompts, eleven hazard categories, and representative attacks, the unsafe-response rate increased from 11.08% at baseline to 18.65% under jailbreak conditions, with an average Resilience Gap of 7.57%. The result quantifies the change in safety performance as the difference between paired baseline and adversarial conditions (the Resilience Gap) rather than reporting only a single-point attack success rate, giving a comparable measure across systems and attacks. The abstract reports explicit sample sizes and an overall statistic; per-system and per-attack confidence intervals or significance tests are not provided.

Accessible systems showed a larger mean Resilience Gap, and attack effectiveness varied substantially across attack categories and hazards. It brings system accessibility together with attack category and hazard type as comparison dimensions, indicating that jailbreak risk is not uniformly distributed across conditions. A qualitative conclusion at the abstract level; specific values and subgroup sample sizes for each system or category are not given.

The benchmark also examines evaluator reliability and sources of measurement error, providing a reproducible methodological foundation for comparative jailbreak evaluation and supporting future expansion across systems, attacks, hazards, and evaluation methods. It treats evaluator calibration and error analysis as components of the benchmark rather than only reporting model scores, bringing measurement reliability into the scope of jailbreak evaluation. The abstract states that the benchmark includes evaluator reliability analysis but does not list specific error metrics or calibration results.

Perspective

The benchmark targets single-turn, text-based jailbreak attacks and is intended for comparing the safety robustness of open-weight large language models under paired baseline and adversarial conditions, using the AILuminate Assessment Standard v1.4 to judge responses. Its methodological framework is designed to expand, so that more systems, attacks, hazard categories, and evaluation methods can be incorporated later, and so that evaluator calibration and measurement-error analysis can become part of routine evaluation. For researchers and engineering teams seeking a reproducible jailbreak evaluation protocol or cross-system comparison of safety performance change, this pipeline offers a directly referenceable starting point.

The abstract does not list specific values per system, per attack, or per hazard category, nor does it state confidence intervals, significance tests, or the level of agreement between human annotation and automated evaluators, so the distribution of the Resilience Gap across conditions remains to be detailed in the full text. Evaluator reliability and sources of measurement error are listed as part of the benchmark's scope, but their specific metrics and conclusions are not presented in the abstract. In addition, this reading is based on the abstract and does not include the body, figures, or appendices, so the details above should be confirmed against the original text.

Sources