#HealthBench
In September, 2024, physicians working with AI did better at the Healthbench doctor benchmark than either AI or physicians alone.

With the release of o3 and GPT-4.1, AI answers are no longer improved on by physicians

Error rates appear to be dropping for newer AI models. openai.com/index/health...
May 13, 2025 at 4:23 AM
Yeah... like, what is the point of boasting about Fable's HealthBench score when you don't actually let it be used on medical tasks?
July 21, 2026 at 4:50 PM
Introducing HealthBench: An evaluation for AI systems and human health openai.com/index/health... #openai #bioinformatics #healthcare
Introducing HealthBench
HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims to provide a shared standard for model perf...
openai.com
May 12, 2025 at 10:51 PM
🏥 HealthBench: el nuevo estándar de IA en salud diseñado por 250+ médicos

https://openai.com/index/healthbench

#IAenSalud #Benchmark #InnovaciónMédica #OpenAI
May 9, 2026 at 1:25 PM
Yesterday OpenAI released HealthBench, an open benchmark of questions for LLMs showing performance on different healthcare tasks. o3, gemini, and grok all do comparably well; what's interesting is that o3 on its own is better than a physician and comparable to a physician with help from o3
May 13, 2025 at 2:39 PM
Trouble brewing
August 7, 2025 at 6:46 PM
Hui, die Kommentare unter dem Artikel.

OpenAI wirbt übrigens jetzt damit, das GPT-5 ein super dupi Gesundheitskompanion ist...
August 15, 2025 at 7:22 AM
MedGemma on HealthBench: Evaluating Open-Source #Medical AI on Consumer Hardware for $39 (preprint) #openscience #PeerReviewMe #PlanP
MedGemma on HealthBench: Evaluating Open-Source #Medical AI on Consumer Hardware for $39
Date Submitted: Mar 21, 2026. Open Peer Review Period: Mar 21, 2026 - Mar 6, 2027.
dlvr.it
March 21, 2026 at 7:02 PM
OpenAI's HealthBench Professional sets a benchmark for evaluating language models in clinician scenarios, enhancing AI reliability and safety. Physician-designed conversations ensure effective AI support for clinical tasks, improving patient care. https://arxiv.org/abs/2604.27470
HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
ArXiv link for HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats
arxiv.org
May 2, 2026 at 5:20 AM
I know today's buzz is all about GPT-5 but this week's earlier announcement on gpt-oss was interesting - the medical evaluations show the open weight models outperform GPT-4o, o1, o3-mini, and o4-mini on HealthBench evals cdn.openai.com/pdf/419b6906...
August 7, 2025 at 8:15 PM
But isn’t AI better than MDs at answering health questions?!

OpenAI says so!
Introducing HealthBench
HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims to provide a shared standard for model perf...
openai.com
July 20, 2025 at 1:42 PM
OpenAi HealthBench: a new benchmark designed to better measure capabilities of AI systems for health. Built with 262 physicians in 60 countries. HealthBench includes 5,000 realistic health conversations each with a custom physician-created rubric to grade model responses.
openai.com/index/health...
Introducing HealthBench
HealthBench is a new evaluation benchmark for AI in healthcare which evaluates models in realistic scenarios. Built with input from 250+ physicians, it aims to provide a shared standard for model perf...
openai.com
May 13, 2025 at 10:17 AM
note how they’re not saying ‘don’t use it in high stakes scenarios’:
September 6, 2025 at 7:37 PM
The updated model matches the performance of OpenAI's most expensive Thinking models on benchmarks like HealthBench and HealthBench Professional, but at a fraction of the cost.
June 19, 2026 at 2:43 AM
OpenAI has introduced HealthBench, a dataset of 5,000 physician-designed health conversations with detailed grading tools, to support standardized evaluation of AI models in health care.
OpenAI releases HealthBench dataset to test AI in health care
OpenAI has unveiled a large dataset to help test how well artificial intelligence (AI) models answer health care questions.
medicalxpress.com
May 14, 2025 at 7:02 AM
HealthBench | Discussion
openai.com
May 12, 2025 at 6:40 PM
HealthBench – An evaluation for AI systems and human health
https://openai.com/index/healthbench/
[comments] [127 points]
May 13, 2025 at 3:03 AM
HealthBench, a New Standard for Evaluating AI in Healthcare Conversations #AI #healthcareAI #genAI #healthcare

thehealthcaretechnologyreport.com/healthbench-...
HealthBench, a New Standard for Evaluating AI in Healthcare Conversations | The Healthcare Technology Report.
thehealthcaretechnologyreport.com
May 29, 2025 at 7:05 PM
OpenAI has made its first big move in health care AI: A big, openly shared benchmark to test how well LLMs work in health care.

Experts I talked to called HealthBench "unprecedented" and "a major step forward," but also raised reasons for caution.

Read more @statnews.com:

🖥️🩺
OpenAI leaps into health care with AI benchmark to evaluate models
OpenAI on Monday released a large set of data for evaluating how well large language models answer questions related to health care. Experts lauded the
www.statnews.com
May 12, 2025 at 6:21 PM
Prof. Mollick, AI's pace is staggering and HealthBench necessary. That said: the narrative of AI > docs is misguided at best, and HealthBench's own methodologies punish brevity (a function of real-world patient encounters). At worst, this narrative fuels unsafe exec staff cuts and staffing ratios.
May 14, 2025 at 9:10 PM
additionally release two HealthBench variations: HealthBench Consensus, which includes 34 particularly important dimensions of model behavior validated via physician consensus, and HealthBench Hard, where the current top score is 32%. We hope that [5/6 of https://arxiv.org/abs/2505.08775v1]
May 14, 2025 at 6:06 AM
Health Bench Explained for Beginners
HealthBench tests AI with real-life medical chats to build safer, smarter healthcare tech.

👀 Watch now! youtube.com/shorts/gs_aJ...

Let’s shape the future of AI in healthcare—engage with this now!

#KashiDHQ #KashiAhmed #AI #ArtificialIntelligence
Health Bench Explained for Beginners
YouTube video by Kashi DHQ
youtube.com
September 13, 2025 at 10:29 AM
OpenAI rolled out new ChatGPT updates including Fast answers, ChatGPT for Clinicians with HealthBench Professional, and ChatGPT for Google Sheets
April 23, 2026 at 8:15 AM