
HealthBench
HealthBench is a testing procedure used to measure how well computer programs answer medical questions. It was released in 2025 by the company OpenAI and is based on evaluation standards from several hundred physicians.
HealthBench is a standardized test for computer programs that answer questions in natural language. It examines a single subject area: health and medicine. The program is presented with realistic conversations, such as the description of symptoms or a follow-up question from a nurse. It is then evaluated on how helpful, safe, and complete the answer was. The standards for this evaluation come from more than 250 physicians from around 60 countries. The test was released in May 2025 by OpenAI, the company behind ChatGPT.
Why medicine needs its own test
Until recently, language programs in medicine were tested mainly with multiple-choice exams. In other words, they were given questions from medical licensing exams, and the correct answers were counted. The best programs quickly reached over 90 percent there. This made the test useless: it could no longer distinguish between good and very good systems.
Above all, a multiple-choice test measures the wrong thing. Nobody types an exam question with four given answers into a chat window. People write incomplete, amateurish, sometimes panicked messages. A good answer must then ask follow-up questions, recognize warning signs, and point out its own limitations. This is exactly what cannot be tested with check marks.
Then there is the risk. A program that botches a movie recommendation does no harm. A program that misses signs of a stroke does. That’s why HealthBench measures not only whether an answer is correct, but also whether it refers the person to a doctor at the right time.
How physicians' opinions become a score
The test consists of 5,000 conversations. They are fictional but built to be realistic and cover emergencies, everyday questions, global health topics, and follow-up questions from healthcare professionals. For each individual conversation, physicians have written their own checklist. In total, that amounts to around 48,000 individual criteria.
A criterion might read: “The response recommends calling emergency services immediately.” Each criterion has a weight in points, both positive and negative. Anyone who forgets an important warning loses points. The same applies to anyone who uses unnecessary technical jargon or states things that are not part of the question. In the end, there is a percentage value: points achieved divided by the maximum possible.
The checking is not done by a human, but by another AI model. This approach is called “model as judge.” It is the weakest point of the procedure, because the judge itself can be wrong. OpenAI therefore checked how often the judging model agrees with human experts. Alongside the main test, there is the variant HealthBench Hard with 1,000 particularly difficult cases. There, the results were initially close to zero percent.
HealthBench in product announcements and stock market news
You will mostly encounter the name in press releases. When a company presents a new model, the table often includes a HealthBench score. It serves as evidence that the system is suitable for the health sector. Such figures are also of interest to investors, because medicine is a huge market for AI providers.
An important distinction must be made here. A good test score is not an approval as a medical device. In Europe, that is decided by regulatory authorities under medical device law, not by a company’s benchmark. The fact that a model performs strongly here does not mean it may be used in a clinic.
A second common misconception concerns the origin. HealthBench comes from OpenAI itself, that is, from a company that sells its own models. The test is publicly available and is also used by others. Nevertheless, the following holds true: whoever designs the exam has an advantage. As with all benchmarks, it is worth comparing multiple sources.