c-84, sector 65, Noida
c-84, sector 65, Noida

Most healthcare leaders will never build a mental health chatbot. MentalHealthBench matters to them for a broader reason. OpenAI’s release names an expectation that now applies to any AI system a patient, a member, or an employee might bring a stressful or health adjacent conversation to. That system will get tested against realistic, unscripted conversations, and it will get tested by clinicians, journalists, regulators, and users long before a procurement team asks the vendor the right question.
This piece is for healthcare founders, provider and payer innovation leaders, digital health operators, and technically fluent executives who sit somewhere between engineering and the board. It walks through what OpenAI actually measured, what independent researchers say a benchmark score cannot show, and the questions a buyer should ask before a patient facing AI product goes live. The questions below sit at the center of healthcare AI governance work today, whether inside a health system or across an AI vendor’s own product line.
A member messages your portal at 2am. The first line is ordinary, something about not sleeping and work being difficult. Your assistant replies warmly and asks a reasonable follow up question. Read on its own, that exchange may look acceptable.
The conversation keeps going. Across the next 20 messages the tone stays supportive, because supportive is what the assistant was built to be. The member describes a belief about a colleague that grows less plausible with each message. The assistant reflects it back each time and never questions it. Any one reply could still look reasonable in isolation. The pattern across all 20 is what creates the risk.
This illustrates the cumulative failure mode the Oxford team measured, and it is the kind of pattern a single response score is not built to see. It is also the version that reaches an incident review, because the transcript can read as a series of reasonable replies until somebody reads it end to end.
Nothing in that scene requires a mental health product. It requires an assistant, a text box, and a member awake at 2am.
The model in that conversation was built by a frontier lab. The safety score was published by the same lab. The conversation happened inside your portal, under your logo, with your member.
When the transcript surfaces, the questions arrive at your organization. A journalist asks your communications team what safeguards were in place. A regulator asks your compliance officer what testing was done before deployment. A family asks your clinical lead why the system answered the way it did.
A vendor’s published score does not answer any of those 3 questions. In MentalHealthBench, OpenAI designed and administered the benchmark, clinicians wrote the rubrics, and GPT-5.6 Sol applied those rubrics to model responses. That makes the benchmark useful and reproducible, but it does not independently verify your configured product, your workflow, or your live deployment. The decisions that belong to you are the decision to deploy, the scope you allow, the handoff you build behind it and the monitoring you run after launch. The evidence that you made those decisions carefully has to be yours as well.
A published benchmark score is a statement about a model. A deployment decision is a statement about your organization, and only you can make that record.

Fig 1 – 1,215 scored conversations against more than 1 billion weekly users
AI mental health chatbot use is no longer a niche behavior. More than 1 billion people use ChatGPT each week, a figure OpenAI states on its September 2026 MentalHealthBench release page. Harvard Medicine Magazine, publishing under Harvard Medical School’s own masthead in May 2026, put a sharper number on what that scale means in practice.
| Signal | Figure |
|---|---|
| ChatGPT weekly users | More than 1 billion |
| Weekly users whose conversations carry explicit suicide planning language | More than 1 million |
| US adults using AI chatbots monthly for health information | 1 in 6 |
| US adults likely to use an AI mental health chatbot within 6 months | 12 percent |
Table 1 – Adoption and usage signals
Those figures reframe what a benchmark release means for healthcare AI governance across an entire portfolio of patient facing systems.
A benchmark score stops being a research footnote once the population behind it outnumbers the patient panel of an entire health system.
The American Psychological Association made the same point in its November 2025 Health Advisory, warning that chatbots and wellness applications are already answering people in crisis without the scientific evidence or regulatory oversight that would normally govern a clinical tool at that volume.
A second scale fact matters as much as the first. The conversations people bring to these systems are rarely obvious emergencies. OpenAI’s own framing of MentalHealthBench states that most existing mental health benchmarks concentrate on emergency scenarios, while the conversations users actually initiate span everyday stress, relationship strain, and ambiguous distress long before anything resembling a crisis classifier would fire. A safety program built only to catch the obvious case misses the volume where most of the exposure sits.

Fig 2 – MentalHealthBench scores 10 behaviors against clinician written rubrics
OpenAI built MentalHealthBench with a deliberately heavy clinical process behind it.
The benchmark’s most disciplined feature is procedural. Two clinicians write every rubric, and a third settles what they cannot agree on.
Every response gets scored across 10 behaviors, from context seeking and clinical reasoning to reality testing calibration and preserving user agency. At release, task clipped scores put GPT-6 Astra at 57.3 percent and Gemini 2.5 Pro at 29.5 percent. A gap this size between model families is exactly the kind of evidence a model selection framework needs to account for once the use case involves patient facing risk.
The executive reading of the score. The 57.3 percent result is useful for comparing models on this benchmark. It is not a 57.3 percent safety rating, a probability that the model is safe, or approval to deploy. Those decisions require evidence about the configured product, full conversations, handoffs, and live monitoring.
How the scoring works. Each rubric criterion carries a point value and some of those values are negative, so a response that does the wrong thing loses points. The signed score adds up every criterion a response meets, penalties included, and divides by the total positive points available for that conversation. That number can fall below zero. The task clipped score, which is the figure OpenAI reports, floors each response at zero so every response contributes between 0 and 1.
Who does the grading. The rubrics are written by clinicians. The grading is not. OpenAI samples 4 responses per conversation and grades each one with GPT-5.6 Sol at high reasoning effort. The benchmark is a model applying criteria that clinicians wrote.
What 57.3 percent is measured against. OpenAI also tested rubric aware completions: responses generated after the model received both the conversation prefix and the grading rubric. Those completions reached 99.0 percent. OpenAI presents that result as a sanity check and a practical near saturation reference. It does not establish a theoretical ceiling. The result shows that the scoring system can recognize an answer deliberately optimized for its rubric. It does not mean a 57.3 percent model is safe in 57.3 percent of conversations or that it satisfied exactly 57.3 percent of clinicians’ expectations.
Clinicians scored below the leading models. Expert written completions reached 38.5 percent, below GPT-6 Astra at 57.3 percent. OpenAI attributes this to clinicians writing conservatively brief responses against a rubric that also rewards the longer, more elaborated answers models tend to produce. That is worth sitting with, because the benchmark rewards a style of answer the clinicians who designed it did not themselves produce.
| Metric | Figure |
|---|---|
| Clinicians involved | 80 plus, across 22 countries |
| Conversations evaluated | 1,215 |
| Behaviors scored | 10 |
| Rubric criteria authored | 5,262 |
| GPT-6 Astra task clipped score | 57.3 percent |
| Gemini 2.5 Pro task clipped score | 29.5 percent |
| Expert authored completions score | 38.5 percent |
Table 2 – MentalHealthBench by the numbers
OpenAI names its own limitations plainly. The benchmark scores the next model response against a supplied conversation prefix, which is a sequence of alternating user and assistant turns that always ends on a user turn. The model is given the whole prefix and writes the next reply, and only that reply is graded. Each score therefore reflects one exchange inside a longer future conversation the benchmark does not itself generate. The user study behind the benchmark, 44 adults across 16 countries, covered only nonacute scenarios and found that expert and user preferences agreed 51.5 percent of the time, compared with 63.4 percent agreement between experts and 62.0 percent between users. OpenAI highlights urgency calibration and reality testing among the weakest behaviors, both directly tied to how a system handles a person in a moment of crisis.
A single response evaluation can confirm a model handles one exchange well. It cannot confirm the model still handles the tenth exchange in the same conversation.
What this means for a buyer. MentalHealthBench can help you compare how models answer the next message under a defined test. It cannot tell you whether your product will stay reliable through a full live conversation, recognize when to hand off, or produce acceptable outcomes after launch. Those are separate tests your deployment still needs.

Fig 3 – Risk accumulates across a conversation, past what one response shows
Four separate bodies of independent research complicate a single benchmark score.
A Nature Medicine study published in August 2026 by researchers at the University of Oxford introduced SIM-VAIL, an adversarial testing framework that ran 810 multi turn conversations across 9 frontier chatbots, including Claude, ChatGPT, Gemini, Grok, and Llama variants. The study found concerning behavior widespread across every chatbot tested, reduced in newer models but not eliminated, with risk building progressively as a conversation continues across multiple turns. Supportive sounding responses could reinforce the psychological mechanisms behind a user’s vulnerability, an effect the authors term a vulnerability amplifying interaction loop, with psychosis and mania producing the highest concerning behavior scores. The authors argue the full conversation is the correct unit of mental health safety measurement. MentalHealthBench scores the next response inside a supplied conversation prefix. The Nature Medicine finding shows why that design can miss an important slice of the risk picture.
Stanford HAI reported in July 2026 that when 3 board certified psychiatrists rated 360 synthetic mental health prompts, their judgments frequently diverged, and averaging the scores produced a number that matched no expert’s actual judgment. A follow up poll of more than 100 psychiatrists at the American Psychological Association’s annual meeting came back almost evenly split on similar cases. Clinicians disagree because they apply different frameworks, safety first, engagement centered, and culturally informed among them, and no amount of averaging reconciles frameworks that start from different premises.
A separate peer reviewed benchmark, PsychiatryBench, published in npj Digital Medicine in April 2026, tested 15 leading models across 5,188 expert annotated clinical items. Even the top performer, GPT-5 Medium at 84.5 percent, showed its weakest results on classifying specific psychiatric disorders and on multi turn follow up and management tasks, the exact areas a procurement conversation is least likely to probe if it stops at one vendor’s headline number.
RAND researcher Ryan McBain wrote in August 2026 that OpenAI released safety scores for its teen focused ChatGPT product without disclosing the prompts, the case counts, or the judging instructions behind them, and without publishing what share of actual teenagers the system correctly identifies. Stanford HAI’s policy team separately counted more than 140 state bills that United States legislatures have introduced on mental health AI, with no comparable federal structure for sharing real world safety data between vendors, independent researchers, and regulators. The Lancet Psychiatry, in a September 2025 viewpoint by clinicians at 3 academic medical centers, called for independent researcher led clinical trials and de identified data sharing as the standard the field has not yet met.
| Study | Source | Key finding |
|---|---|---|
| SIM-VAIL | Nature Medicine, University of Oxford, Aug 2026 | Concerning behavior builds progressively across multi turn conversations |
| Rater disagreement study | Stanford HAI, Jul 2026 | Board certified psychiatrists frequently disagree, and averaging their scores produces a number that matches no individual expert’s judgment |
| PsychiatryBench | npj Digital Medicine, Apr 2026 | A high score on one benchmark does not carry over to weaker areas such as multi turn follow up and management |
| Independent verification commentary | RAND, Stanford HAI policy team, The Lancet Psychiatry | No established path yet exists for outside researchers or regulators to check vendor reported safety data |
Table 3 – What each additional study adds
Turn the research into 3 separate decisions. Use the benchmark to compare models. Use multi turn testing to evaluate conversation level behavior. Use independent, product specific validation and post launch monitoring to decide whether your deployment is ready. One result cannot substitute for the other 2.
A score without an audit trail is a claim. Independent verification is what turns a claim into evidence.
Clixlogix has raised a version of this same caution in AI powered medical imaging, where a strong benchmark score still needs local, site specific validation before it earns clinical trust. The same expectation holds across every clinical AI category, a published number opens the conversation and does not close it.

Fig 4 – The 7 questions a healthcare buyer asks before deployment
A benchmark score answers a narrower question than most procurement conversations assume. AI governance in healthcare depends on documentation a vendor can produce on demand, well beyond a single published number. The following 7 questions turn the findings above into a due diligence checklist for any AI system that will talk to a patient, a member, or an employee about a stressful or health adjacent topic.
A question is only useful if an answer changes what you do. The table below pairs each question with the answer a prepared vendor gives, and with the action to take when that answer does not arrive.
| # | Ask | A good answer sounds like | If they cannot answer |
|---|---|---|---|
| 1 | Which behaviors are tested, and can we see each result? | A per behavior table, with the weakest 2 named without prompting | Treat the aggregate as unverified and require the breakdown before signature |
| 2 | Who wrote the criteria, and how was disagreement resolved? | Named clinical roles, a written adjudication process, criteria available for review | The score measures agreement with an undisclosed standard. Commission your own rubric for your top 20 scenarios |
| 3 | Did the evaluation test full conversations or only the next reply? | They generate and score full conversations, and can show how scores move between turn 1 and turn 20 | Assume multi turn risk is unmeasured. Restrict the assistant to short scoped exchanges with a hard handoff |
| 4 | Who graded the responses, and how often did the graders agree? | A published agreement figure, and a clear statement of whether graders were human or model | Treat any single score as a point estimate with unknown spread and do not use it to rank vendors |
| 5 | Can an independent party reproduce or verify the result? | A path for an outside researcher or regulator, or published prompts and case counts | Budget for your own pre deployment evaluation, because nothing external will catch a regression |
| 6 | How is human handoff triggered, and what is its measured miss rate? | A measured handoff rate on a defined trigger set, with the false negative rate | Build the handoff yourself at the application layer and do not rely on the model to initiate it |
| 7 | What is monitored after launch, and who owns escalation? | Live monitoring on sampled conversations, a defined escalation path, a named owner | You are running an unmonitored clinically adjacent system. Add sampling and review before launch |
Table 4 – What each answer should trigger
Healthcare AI compliance increasingly means producing this evidence before a system reaches production. A vendor with credible answers to all 7 has done work most of the market has not yet made visible.
A strong score with no supporting detail is a demonstration. A strong score with an audit trail is a safety program.
The regulatory position is active and unsettled, which is the expensive combination. The FDA’s Digital Health Advisory Committee met on 6 November 2025 to weigh regulation of generative AI enabled digital mental health medical devices. Stanford HAI counts more than 140 state bills introduced across the United States on mental health AI, while federal oversight remains fragmented and has not produced a comparable shared framework for safety data. The American Psychological Association issued a Health Advisory in November 2025 warning that these tools are already answering people in crisis without the oversight a clinical tool would normally carry.
A team that runs the checklist before deployment produces a dated record of what was tested, by whom, against which criteria, and what was done about the gaps. That record is what a regulator, a journalist or a claims process asks for, and it is not something that can be assembled after the fact.
The 7 questions take one meeting. After an incident the same questions still get asked, by somebody else and on their timetable.
Clixlogix works with healthcare and health adjacent SMBs across the United States and internationally, building the AI systems that sit inside patient portals, member communications, and clinical workflow tools. The questions in the checklist above are the same questions our delivery teams raise before a client’s AI system goes live, whether a team builds that system in house, licenses it from a vendor, or assembles it from several. An architecture review against this checklist, completed before deployment, is typically the fastest way for a healthcare leader to know where a system actually stands, well before an incident forces the question.
A structured evaluation against this checklist looks at 7 evidence areas.
Teams weighing an AI vendor or an internal build in this space can talk to our team about a structured healthcare AI governance review.
No. OpenAI’s own release states that no benchmark captures everything that matters in a personal conversation. MentalHealthBench scores the next response from a supplied conversation prefix, while several of the safety concerns documented in peer reviewed research, including risk that accumulates across a full conversation, sit outside what that design can show.
No credible source in this record makes that claim, and several explicitly warn against it. The American Psychological Association’s Health Advisory states plainly that AI chatbots should not substitute for licensed mental health professionals, and The Lancet Psychiatry frames the open research question narrowly, as whether these tools can responsibly support access gaps.
Independent verification. RAND, Stanford HAI, and the American Psychological Association each point to the same gap, real world safety data currently sits inside the vendor with no established path for independent researchers or regulators to check it. A vendor willing to open that data to outside review has cleared a bar most of the market has not.
Unevenly. MentalHealthBench measures harm avoidance and user agency as separate behaviors but does not quantify how often a system successfully hands a conversation to a human. Multiple independent commentators flagged this as a gap during the benchmark’s release, and it remains an open measurement question industry wide.
Active and unsettled. The FDA’s Digital Health Advisory Committee met on 6 November 2025 to weigh regulation of generative AI enabled digital mental health medical devices, and Stanford HAI counts more than 140 state bills addressing mental health AI introduced across the United States, with no comparable federal framework yet in place.
MentalHealthBench is a genuine step forward, built with more real clinical input than most benchmarks in this space have carried before it. OpenAI designed and administered the evaluation, clinicians authored its rubrics, and an OpenAI model applied them to the next response from each supplied conversation prefix. What remains unverified is whether any buyer’s configured product stays safe across full live conversations and real deployment conditions at the scale AI mental health use now demands.
A benchmark score is a starting signal. The checklist above is what turns it into a governance program.
A healthcare leader deploying a patient facing AI system does not need to wait for that verification to catch up. Responsible AI in healthcare turns a benchmark score into the starting point for the checklist above. The leaders who use it before deployment will spend less time explaining an incident after the fact.

Pushker is the founder of Clixlogix. Give him a messy operation and he finds the leverage point, then builds the fix himself. He works at the edge of what AI can actually do inside a business, and writes about what he finds there.
We are here to answer your questions 24/7
