MentalHealthBench tests AI responses to 1,215 synthetic mental health conversations using criteria developed by more than 80 licensed experts. Its scenarios span everyday concerns and emergencies, so the test asks more than whether a model avoids an unsafe reply.
Published byEdited with GPT-6 SolImages generated with GPT Image 2
Research answer

Create a landscape editorial hero image for this Studio Global article: What is OpenAI’s MentalHealthBench, why was it created, how were its conversations and expert scoring criteria developed and evaluated, what. Article summary: MentalHealthBench is OpenAI’s open benchmark for testing whether AI gives helpful, safe responses across 1,215 realistic mental-health conversations. OpenAI created it because earlier evaluations focused heavily on crise. Topic tags: general, general web. Style: premium digital editorial illustration, source-backed research mood, clean composition, high detail, modern web publication hero. Use reference image context only for broad subject, composition, and topical grounding; do not copy the exact image. Avoid: logos, brand marks, copyrighted characters, real person likenesses, fake screenshots, UI text, readable text, watermarks, charts with fake numbers, clic
OpenAI’s MentalHealthBench evaluates whether an AI response is helpful and safe for the particular mental-health conversation in front of it. Released as an open benchmark, it pairs 1,215 synthetic conversations with situation-specific criteria written by licensed mental-health experts. OpenAI created it to address a gap in evaluations that concentrated on emergencies and broad safety rules rather than the wider range of support people may seek. 1
3
The conversations are synthetic but designed to represent varied situations encountered in mental-health discussions. Some provide earlier context, allowing the evaluation to test whether a model takes relevant details into account when answering the final user message. The set is a collection of test cases, not an estimate of how frequently each situation occurs in ChatGPT. 1
More than 80 licensed mental-health experts from 22 countries helped develop the benchmark. Experts reviewed each conversation and wrote criteria for assessing a response to its final message. Each conversation received review from at least three experts; OpenAI says a criterion entered the final benchmark only when all reviewers agreed to it. 1
4
The released set contains 5,262 expert-written criteria. Each targets a specific response behavior: positive weights reward helpful actions, while negative weights penalize harmful ones. An automated grader then assesses a model’s response against the criteria individually. This makes the score a measure of alignment with those rubrics—not a direct measure of a person’s clinical outcome. 1
4
MentalHealthBench ranges from everyday, non-acute concerns to serious distress and urgent emergencies. Its scenarios include adults, teens, caregivers and clinicians, with variation across languages and regions. The criteria can examine whether a response seeks needed context, respects a person’s autonomy, offers appropriate practical guidance and responds proportionately to risk. 1
4
OpenAI reports improvement among the AI systems it tested, alongside persistent difficulties with asking for appropriate context and calibrating urgency. The supplied results do not establish a reliable model-by-model ranking for this article, and better benchmark performance should not be read as a guarantee of safe behavior in every conversation. 1
OpenAI also considered feedback from real ChatGPT users. That perspective can complement clinicians’ judgments, but it is not interchangeable with clinical and safety criteria. The available source material does not establish which specific priorities users and clinicians ranked differently, so it would be misleading to claim a particular divide. 1
The benchmark measures responses; it is not itself a crisis-support feature. Separately, OpenAI says it has expanded access to crisis hotlines, redirected some sensitive conversations to safer models and added reminders to take breaks during long sessions. Those measures provide context for its safety work, but they do not make ChatGPT a substitute for therapy, professional care or real-world help in an emergency. 11
15
Studio Global AI
This page includes a source-backed answer you can continue inside Studio Global.
MentalHealthBench tests AI responses to 1,215 synthetic mental health conversations using criteria developed by more than 80 licensed experts.
MentalHealthBench tests AI responses to 1,215 synthetic mental health conversations using criteria developed by more than 80 licensed experts. Its scenarios span everyday concerns and emergencies, so the test asks more than whether a model avoids an unsafe reply.