OpenAI launches benchmark for AI in mental health conversations
AI firm OpenAI has released MentalHealthBench, an open benchmark designed to test how AI systems respond to realistic mental health conversations.

Artificial intelligence firm OpenAI announced on September 23 the release of MentalHealthBench, an open benchmark for evaluating AI systems' performance in mental health conversations. The tool aims to assess how AI responds to various scenarios, from non-acute discussions to emergencies.
OpenAI stated that the benchmark was co-created with over 80 licensed mental health professionals from 22 countries. It utilizes synthetic chat data covering a range of situations, including adult and teen conversations, caregiver interactions, and clinical settings, across multiple languages and regions.
The company used an automated grader, GPT-5.6 Sol, to score AI model responses against expert-defined criteria. A parallel study involving 44 adults from 16 countries who had used AI for emotional support revealed differing priorities. Users valued tone and actionable advice more, while experts emphasized context gathering and careful interpretation of ambiguity.
This initiative follows OpenAI's recent efforts to strengthen AI responses in sensitive conversations and introduce features like Trusted Contact in ChatGPT. The launch also occurs amid broader regulatory attention to digital mental health tools and legal challenges concerning AI's impact.