OpenAI’s MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations

OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

OpenAI’s MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations

OpenAI built a rigorous new benchmark with 80-plus mental health experts to test how well AI handles sensitive conversations, and its newest model leads the pack

Mental health conversations are some of the most consequential interactions a person can have. OpenAI has now built a formal measuring stick for exactly this problem, and its latest model is the first to be tested against it.

OpenAI released MentalHealthBench on September 23, 2026, an open benchmark designed to evaluate how AI models perform in realistic mental health conversations. GPT-6 Astra, the company’s most recent flagship model, scored 57.3 on the benchmark, placing it ahead of its predecessor GPT-4o.

What MentalHealthBench actually measures

OpenAI developed the benchmark by working with more than 80 licensed mental health professionals across 22 countries. The scoring rubric runs on a weighted scale from -1 to +10 for each evaluated response. Scores are determined by four primary criteria: safety, contextual relevance, preservation of user agency, and the ability to provide actionable guidance. A response that steers a user toward harmful behavior can score negative; a response that correctly identifies an emergency and provides appropriate escalation paths scores near the top.

Advertisement

The benchmark evaluates AI across distinct user personas rather than treating all users as interchangeable. The tested personas include adults, teens between the ages of 13 and 17, caregivers, and clinicians.

The synthetic conversations used in the evaluation were built to mirror real ChatGPT usage patterns. OpenAI has released the benchmark’s methodology and synthetic data publicly, which allows independent researchers to run their own evaluations and challenge OpenAI’s findings.

GPT-6 Astra’s score in context

GPT-6 Astra launched on September 3, 2026, roughly three weeks before the benchmark’s public release.

Dr. Arthur Evans, CEO of the American Psychological Association, pointed to the importance of addressing the full range of mental health needs, from everyday emotional support to acute crisis intervention.

Why this benchmark matters beyond OpenAI

Until now, there has been no widely accepted, expert-validated standard for evaluating AI in mental health contexts. Other AI developers, including Anthropic, Google DeepMind, and Meta, can now run their own models against the same rubric and publish comparable results. Academics, clinicians, and policy researchers can also use the publicly released methodology and synthetic data to conduct independent work on AI and mental health safety.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
OpenAI’s MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations
OpenAI’s MentalHealthBench rates GPT-6 Astra at 57.3 for mental health conversations

OpenAI built a rigorous new benchmark with 80-plus mental health experts to test how well AI handles sensitive conversations, and its newest model leads the pack

OpenAI official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Mental health conversations are some of the most consequential interactions a person can have. OpenAI has now built a formal measuring stick for exactly this problem, and its latest model is the first to be tested against it.

OpenAI released MentalHealthBench on September 23, 2026, an open benchmark designed to evaluate how AI models perform in realistic mental health conversations. GPT-6 Astra, the company’s most recent flagship model, scored 57.3 on the benchmark, placing it ahead of its predecessor GPT-4o.

What MentalHealthBench actually measures

OpenAI developed the benchmark by working with more than 80 licensed mental health professionals across 22 countries. The scoring rubric runs on a weighted scale from -1 to +10 for each evaluated response. Scores are determined by four primary criteria: safety, contextual relevance, preservation of user agency, and the ability to provide actionable guidance. A response that steers a user toward harmful behavior can score negative; a response that correctly identifies an emergency and provides appropriate escalation paths scores near the top.

Advertisement

The benchmark evaluates AI across distinct user personas rather than treating all users as interchangeable. The tested personas include adults, teens between the ages of 13 and 17, caregivers, and clinicians.

The synthetic conversations used in the evaluation were built to mirror real ChatGPT usage patterns. OpenAI has released the benchmark’s methodology and synthetic data publicly, which allows independent researchers to run their own evaluations and challenge OpenAI’s findings.

GPT-6 Astra’s score in context

GPT-6 Astra launched on September 3, 2026, roughly three weeks before the benchmark’s public release.

Dr. Arthur Evans, CEO of the American Psychological Association, pointed to the importance of addressing the full range of mental health needs, from everyday emotional support to acute crisis intervention.

Why this benchmark matters beyond OpenAI

Until now, there has been no widely accepted, expert-validated standard for evaluating AI in mental health contexts. Other AI developers, including Anthropic, Google DeepMind, and Meta, can now run their own models against the same rubric and publish comparable results. Academics, clinicians, and policy researchers can also use the publicly released methodology and synthetic data to conduct independent work on AI and mental health safety.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.