OpenAI launches MentalHealthBench to evaluate AI in mental health conversations

OpenAI introduced MentalHealthBench, an open benchmark designed to measure how AI systems respond in realistic mental health conversations. Published September 23, 2026, the benchmark was co-created with more than 80 licensed psychologists and psychiatrists from 22 countries and aims to evaluate model behavior across a wide range of scenarios and user personas.

What the benchmark measures

MentalHealthBench assesses model capabilities on behaviors experts identified as important for mental health interactions: safety, seeking context, preserving user agency, and offering actionable guidance when appropriate. Scenarios span a full acuity spectrum—non-acute (53.5%), high-acuity (18.2%), and emergent (28.3%)—and cover adults (68.1%), teens aged 13–17 (21.2%), clinicians (5.8%), and caregivers (4.9%). For the teen persona, the age range was specified to models through a system message and cases were reviewed by clinicians with youth mental health expertise.

How it was built

Using privacy-preserving methods, OpenAI generated synthetic conversations reflecting real-world use. Each synthetic conversation was reviewed by at least three experts, who produced weighted rubric criteria for evaluating model responses. Criteria carry weights from -10 to +10 to reward beneficial behaviors and penalize harmful ones; items were retained only if agreed on by at least two experts and not contradicted by a third.

OpenAI used an automated grader, GPT‑5.6 Sol, to score model responses against the expert-written rubrics. The company says this process and the evaluation settings are described in an accompanying paper.

Findings, user perspectives and availability

OpenAI reports that results on MentalHealthBench show steady improvement in recent frontier models and that finer-grained measurements reveal different strengths across behavior dimensions—for example, increased ability to seek context in more advanced models. The release also includes a comparison between experts’ criteria and feedback from 44 adults across 16 countries and 14 languages, whose ratings emphasized practical next steps and tone; their review was limited to non-acute conversations.

OpenAI cautioned that ChatGPT is not a substitute for therapy or professional care, and highlighted related product changes such as expanded crisis resources, Trusted Contact, and ChatGPT for Teens. MentalHealthBench is released openly so researchers and developers can examine methods, run evaluations, and build on the benchmark. MentalHealthBench was published September 23, 2026 and is available for others to use.


Original source: OpenAI News

Leave a Comment