Skip to main content
Back to timeline
OpenAISource publication:

MentalHealthBench: An Open Benchmark for Realistic Mental Health Conversations

Synopsis

This work co-created MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries, using privacy-preserving techniques to generate synthetic mental health conversations that span non-acute, high-acuity, and emergency situations and four user personas (adults, teens aged 13-17, caregivers, and clinicians), with experts writing rubric criteria weighted from -10 to +10 (each conversation reviewed by at least three experts, retaining only criteria agreed by at least two and not contradicted by a third) and an automated grader, GPT-5.

AI-generated editorial illustration: Introducing MentalHealthBench

Interpretation

It introduces and openly releases MentalHealthBench, extending evaluation of models in mental health settings beyond a primary focus on emergency scenarios and broad predefined criteria to the full spectrum of non-acute, high-acuity, and emergency conversations. The text states that most prior evaluations focused primarily on emergency scenarios and measured success with broad, predefined criteria; this benchmark instead writes per-conversation expert criteria, so it can assess whether responses align with expert guidance for each situation rather than only whether disallowed responses were avoided. The benchmark was co-created with more than 80 licensed psychologists and psychiatrists across 22 countries, speaking 19 languages and representing nearly 20 mental health subspecialties, with each conversation reviewed by at least three experts.

It establishes an interpretable, multi-faceted scoring system in which experts write single-aspect criteria for the last user message, weighted from -10 to +10, where positive points reward beneficial behaviors and negative points penalize harmful ones, and larger values indicate greater clinical importance. Compared with broad success judgments, this design breaks behaviors such as asking the right question or providing the best possible advice into separate weighted items, allowing the overall score to be decomposed into ten dimensions of mental health behavior. Only criteria agreed upon by at least two experts and not contradicted by a third were retained, so the final rubrics reflect shared expert judgment; grading is performed by the automated grader GPT-5.6 Sol against the expert-written criteria.

Evaluating a wide range of models on the benchmark shows steady improvement of AI systems in helping people navigate mental health situations, and the ability to seek context appropriately increased with more advanced models. The text attributes these improvements to investments by OpenAI and other model providers in improving how models navigate mental health conversations, and notes that models with similar overall scores can have different strengths across the ten behaviors. Results are based on evaluation over the entire dataset, measuring whether a model's response demonstrates all the ideal behaviors experts identified for each scenario while avoiding less desirable behavior; the text does not report specific scores or a model list.

A separate analysis with 44 adults who had used AI for mental health or emotional support, representing 16 countries and 14 languages, compared expert guidance with what people find helpful. User perspectives highlighted practical next steps and tone, which were less emphasized in expert guidance, while experts placed greater emphasis on gathering relevant context and carefully interpreting ambiguous situations; this analysis did not change the benchmark's final expert-consensus scoring criteria. Participant review was limited to non-acute conversations to avoid exposing them to potentially distressing high-acuity material, and the text describes this as a separate analysis alongside the benchmark.

Perspective

The benchmark targets four user types (adults, teens aged 13-17, caregivers, and clinicians) and covers non-acute, high-acuity, and emergency situations across multiple languages and regions; its scoring is based on expert consensus and measures whether a model's response demonstrates the ideal behaviors experts identified for each scenario while avoiding less desirable behavior. The text explicitly states that ChatGPT is not a substitute for therapy or professional care, and that the benchmark's purpose is to guide people toward real-world support such as localized crisis hotlines or someone they trust. The teen persona is indicated through a system message stating the user is between 13-17; this approach is designed to work across model providers, though the text notes it may not capture all safeguards built into individual products.

The text does not report specific model scores, rankings, or the list of evaluated models, nor does it disclose the number of synthetic conversations or details of language distribution, so the magnitude of improvement cannot be judged from this article. The user-perspective analysis covers only non-acute conversations, leaving open whether user and expert judgments align in high-acuity and emergency situations. In addition, scoring relies on the automated grader GPT-5.6 Sol, so how closely it matches expert human judgment, and how comparable the benchmark is across model providers with differing built-in product safeguards, remain directions a careful reader would continue to watch.

Sources