Capital & Compute
Benchmark· Domain & professional work· Checked 2026-07-27

HealthBench

HealthBench measures the quality and safety of open-ended clinical conversations rather than performance on medical multiple-choice exams. Responses are graded against rubrics written by practising physicians. It became the 2026 reference for medical AI because it tests the thing that actually matters clinically: what the model says to a person, not whether it can pass a test.

Key facts about the HealthBench benchmark
What it measuresOpen-ended clinical conversation quality and safety, graded against rubrics written by practising physicians.
Built byOpenAI (Arora, Wei, Soskin Hicks et al.), 2025
Format5,000 multi-turn conversations with users and health professionals, scored against 48,562 rubric criteria written by 262 physicians
Scoring metricRubric score, graded by a model grader against physician-written criteria
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard

How HealthBench works

The benchmark contains 5,000 multi-turn conversations between a model and either an individual user or a healthcare professional. Each conversation has its own rubric, and across the set there are 48,562 unique criteria written by 262 physicians, spanning contexts such as emergencies, transforming clinical data and global health, and behavioural dimensions such as accuracy, instruction following and communication. A model grader scores responses against those criteria.

History and current status

OpenAI released HealthBench in 2025. It arrived against a background of medical benchmarks that had gone stale: MedQA, built from board-exam questions, was saturated and heavily contaminated, and "AI passes the medical licensing exam" headlines had stopped being informative. HealthBench replaced the exam format with rubric-graded conversation, and Stanford MedHELM provides the independent academic counterpart.

What the score does not tell you

Two structural caveats. It was built and is run by a model vendor, which is the same independence problem that dogs vendor-run benchmarks generally. And it is graded by a model against physician-written criteria, so the grader’s own reliability is part of the measurement. The physician-written rubrics are a genuine strength; the vendor-built, model-graded pipeline around them is where to apply scepticism.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports HealthBench, and how to read it

OpenAI reports it for its own models, and it now appears widely in medical AI coverage. For a question about fitness for a specific clinical workflow rather than general conversation quality, Stanford MedHELM is the better citation because it is independent and organised around a clinician-validated task taxonomy.

2025
First released
OpenAI (Arora, Wei, Soskin Hicks et al.)
Active
Status today
As of July 27, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

HealthBench: frequently asked questions

What is HealthBench?
HealthBench is a 2025 OpenAI benchmark of 5,000 multi-turn health conversations, scored against 48,562 rubric criteria written by 262 physicians. Unlike multiple-choice medical benchmarks it evaluates open-ended clinical conversation quality and safety.
How is HealthBench different from MedQA?
MedQA is multiple-choice board-exam questions, now saturated and contaminated. HealthBench is open-ended conversation graded against physician-written rubrics. Passing an exam and safely handling a patient conversation are different skills, and HealthBench measures the second.
Is HealthBench independent?
No. It was built and is run by OpenAI, and responses are graded by a model rather than by the physicians who wrote the rubrics. Stanford MedHELM is the independent academic alternative, organised around a clinician-validated taxonomy of real clinical tasks.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory