We’re demonstrating how frontier models have continued to improve in realistic…
- Source
- OpenAI
- Date
We’re demonstrating how frontier models have continued to improve in realistic mental health conversations with MentalHealthBench. This new open benchmark was built with input from more than 80 mental health clinicians. We’re releasing it openly so other researchers can examine the methods, run their own evaluations, and build on the work.

Most mental health benchmarks focus on emergency situations. MentalHealthBench is designed to cover the full spectrum of mental health conversations that people bring to AI - from everyday support to more acute crisis scenarios.

- eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Gives practitioners a clinician-validated for a high-stakes conversational domain, useful for benchmarking safety and quality of AI mental health responses.
articleA new benchmark for evaluating patient-facing health AI agentswww.amazon.science
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert KirkarticlePiloting the world's first double-blind AI evaluationsdeepmind.google
Checking sign-in…
Loading comments…


