📣 Send us your press release
Site updates every 15 minutes
Technology

OpenAI releases MentalHealthBench to evaluate AI's mental health capabilities

OpenAI has launched MentalHealthBench, an open benchmark comprising 1215 synthetic mental health conversations, designed to assess AI systems' ability to handle real-world scenarios.

24 September 2026
OpenAI releases MentalHealthBench to evaluate AI's mental health capabilities
Image is an AI-generated illustration

Artificial intelligence research company OpenAI has released MentalHealthBench, a new tool intended to evaluate AI models' proficiency in handling conversations related to mental health.

The benchmark includes 1215 synthetic psychological dialogues, covering a wide range of situations from everyday concerns to emergency scenarios. Its development involved over 80 licensed psychologists and psychiatrists from 22 countries, utilizing 19 languages and representing nearly 20 specialized fields within mental health. Each dialogue is accompanied by expert-defined scoring criteria.

OpenAI states that current evaluation methods often focus narrowly on emergencies and employ broad standards. MentalHealthBench aims to address this gap by providing a more comprehensive assessment of AI performance across diverse mental health interactions. The company hopes the public release will allow other researchers to scrutinize the methodology, conduct independent evaluations, and build upon the work.

Initial test results indicate that OpenAI's own models, such as GPT-6 Astra, performed highest on the benchmark. However, the findings also highlighted areas for improvement, particularly in models' ability to gather appropriate background information and accurately gauge urgency. OpenAI emphasizes that ChatGPT is not a substitute for professional therapy or medical services.

The MentalHealthBench dataset is available for download. OpenAI requests researchers refrain from publishing the dataset's content verbatim to prevent it from entering model training data and potentially contaminating future evaluations.

Original source: ithome.com