Product Launches

OpenAI's MentalHealthBench scores GPT-6 Astra at 57.3% in mental health AI testing

Share
OpenAI's MentalHealthBench scores GPT-6 Astra at 57.3% in mental health AI testing

OpenAI has launched MentalHealthBench, an open benchmark for evaluating AI responses in mental health conversations, co-created with 80+ experts across 22 countries. The benchmark scored GPT-6 Astra highest at 57.3%, with notable performance gaps across different models.

TL;DR

  • OpenAI introduces MentalHealthBench to assess AI's mental health conversation capabilities, co-developed with global experts.
  • GPT-6 Astra leads with a 57.3% score, followed by GPT-6 Sol and Claude Opus 5.5, highlighting room for improvement in AI mental health support.
  • The benchmark evaluates clinical accuracy, empathy, urgency recognition, and more, with results broken down by conversation type and user profile.

What happened

OpenAI has released MentalHealthBench, an open benchmark designed to evaluate AI systems' responses in mental health and emotional support conversations. The benchmark was developed with over 80 licensed psychologists and psychiatrists from 22 countries, representing nearly 20 mental health subspecialties and speaking 19 languages. MentalHealthBench assesses more than just safety, evaluating behaviors such as clinical accuracy, context appropriateness, user agency preservation, actionable guidance, empathy, urgency recognition, and harm avoidance.

The benchmark uses synthetic conversations reflecting real-world patterns, covering adults, teens, caregivers, and clinicians across multiple languages and regions. These conversations vary in urgency, with 53.5% being non-acute, 18.2% high-acuity, and 28.3% emergencies. More than half of the conversations contain over five messages, allowing evaluation of AI responses across exchanges, not just isolated prompts.

OpenAI tested models from multiple providers on MentalHealthBench. GPT-6 Astra recorded the highest overall score at 57.3%, followed by GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%. Older models like GPT-4o scored significantly lower at 32.1%, and Gemini 2.5 Pro at 29.5%. The benchmark also allows performance breakdown by conversation urgency, user profile, and individual behavior.

Why it matters

MentalHealthBench provides a standardized way for researchers and developers to evaluate and improve AI's mental health conversation capabilities. By making the benchmark open, OpenAI encourages collaboration and further development in this critical area. The benchmark's comprehensive evaluation criteria, developed by mental health experts, ensure that AI responses are assessed for clinical accuracy, empathy, and other crucial factors.

For AI/ML developers and startup founders, MentalHealthBench offers a valuable tool to test and refine their models' mental health support capabilities. The benchmark's detailed scoring and performance breakdown can help identify areas for improvement and guide development efforts. For tech investors, the benchmark highlights the competitive landscape and the ongoing efforts to enhance AI's role in mental health support.

However, it's important to note that MentalHealthBench scores are not measures of clinical effectiveness but rather evaluations of whether responses demonstrate behaviors identified by expert reviewers. The benchmark is part of a broader effort by OpenAI to advance mental health research and model safety as AI systems continue to evolve.

Key facts

  • MentalHealthBench was co-created with over 80 licensed psychologists and psychiatrists from 22 countries.
  • The benchmark evaluates AI responses across ten areas, including clinical accuracy, empathy, urgency recognition, and harm avoidance.
  • GPT-6 Astra scored highest at 57.3%, followed by GPT-6 Sol at 53.9% and Claude Opus 5.5 at 52.4%.
  • Older models like GPT-4o scored 32.1%, and Gemini 2.5 Pro scored 29.5%.
  • The benchmark uses synthetic conversations covering adults, teens, caregivers, and clinicians across multiple languages and regions.
  • 53.5% of conversations are non-acute, 18.2% are high-acuity, and 28.3% are emergencies.
  • More than half of the conversations contain over five messages, allowing evaluation of AI responses across exchanges.
  • MentalHealthBench is openly available for researchers and developers to inspect, evaluate, and build upon.

Context

The launch of MentalHealthBench comes at a time when AI's role in mental health support is gaining increasing attention. AI systems have the potential to provide accessible and scalable mental health support, but ensuring their safety, accuracy, and empathy is crucial. MentalHealthBench is a significant step towards standardizing the evaluation of AI's mental health conversation capabilities.

OpenAI's initiative aligns with broader industry efforts to enhance AI's role in mental health. Other companies and researchers are also working on developing AI models and tools for mental health support, highlighting the growing recognition of AI's potential in this field. As AI systems continue to evolve, benchmarks like MentalHealthBench will play a vital role in ensuring their effectiveness and safety.

The mental health landscape is complex and diverse, requiring AI systems to be evaluated across a wide spectrum of conversations and user profiles. MentalHealthBench's comprehensive evaluation criteria and global expert involvement reflect this complexity, providing a robust framework for assessing AI's mental health support capabilities.

Topics

Related coverage

Join the discussion

Have a take on this story? Weigh in with our community on Facebook.

💬 Discuss on Facebook →