Public articles linked to the same research event.
OpenAI This work co-created MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries, using privacy-preserving techniques to generate synthetic mental health conversations that span non-acute, high-acuity, and emergency situations and four user personas (adults, teens aged 13-17, caregivers, and clinicians), with experts writing rubric criteria weighted from -10 to +10 (each conversation reviewed by at least three experts, retaining only criteria agreed by at least two and not contradicted by a third) and an automated grader, GPT-5.
This work co-created MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries, using privacy-preserving techniques to generate synthetic mental health conversations that span non-acute, high-acuity, and emergency situations and four user personas (adults, teens aged 13-17, caregivers, and clinicians), with experts writing rubric criteria weighted from -10 to +10 (each conversation reviewed by at least three experts, retaining only criteria agreed by at least two and not contradicted by a third) and an automated grader, GPT-5.
This work co-created MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries, using privacy-preserving techniques to generate synthetic mental health conversations that span non-acute, high-acuity, and emergency situations and four user personas (adults, teens aged 13-17, caregivers, and clinicians), with experts writing rubric criteria weighted from -10 to +10 (each conversation reviewed by at least three experts, retaining only criteria agreed by at least two and not contradicted by a third) and an automated grader, GPT-5.
This work co-created MentalHealthBench, an open benchmark built with more than 80 licensed mental health experts from 22 countries, using privacy-preserving techniques to generate synthetic mental health conversations that span non-acute, high-acuity, and emergency situations and four user personas (adults, teens aged 13-17, caregivers, and clinicians), with experts writing rubric criteria weighted from -10 to +10 (each conversation reviewed by at least three experts, retaining only criteria agreed by at least two and not contradicted by a third) and an automated grader, GPT-5.