GPT-5.5 Instant’s health assessment capabilities have jumped, and doctors’ evaluations surpass those of their human counterparts

GPT-5.5 Instant reaches state-of-the-art performance on HealthBench and HealthBench Professional benchmarks, with answers better than those of human doctors in double-blind physician evaluations.

OpenAI announces GPT-5.5 Instant, a major breakthrough in health assessment. The model reaches state-of-the-art performance on both HealthBench and HealthBench Professional benchmarks. Even more strikingly, in double-blind evaluations conducted by doctors, GPT-5.5 Instant's answers were found to be better than those of human doctors.

HealthBench covers basic knowledge assessment in multiple medical fields, including common departments such as internal medicine, pediatrics, and surgery. HealthBench Professional focuses on more professional medical subfields, testing the model's ability in specialized diagnosis and treatment plan recommendations.

OpenAI said this improvement is due to deep reinforcement learning of medical literature and clinical guidelines during the GPT-5.5 training process. Models are not only capable of answering medical questions but also maintain consistency across complex clinical reasoning tasks.

This capability has been integrated into ChatGPT, allowing users to access health-related information in conversations. But OpenAI emphasizes that this is still an auxiliary tool, not a medical device, and should not be used as a substitute for professional medical advice.

Medical AI is a key vertical area that major model manufacturers compete for. GPT-5.5 Instant's "better than human" performance in double-blind evaluations by doctors is landmark in the industry - this is the first time a general conversation model has crossed this threshold in clinical knowledge question and answer. For the domestic medical AI track, this is both a catch-up goal and a warning: the in-depth capabilities of general large models in professional fields are rapidly breaking through the advantage areas of disease-specific models.

Copyright: Content sourced from OpenAI Official Blog . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...