ChatGPT improves its free health answers
OpenAI says GPT-5.5 Instant, free in ChatGPT, now matches its frontier models on its hardest health evals, which are scored with physician-written rubrics.
OpenAI said on June 18 that GPT-5.5 Instant, the model behind the free tier of ChatGPT, now scores about the same as its frontier Thinking models on the company’s toughest health evaluations. That matters because the cheaper, faster model is what most people actually hit: OpenAI says more than 230 million people a week ask ChatGPT health and wellness questions.
The scoring leans on two in-house benchmarks, HealthBench and HealthBench Professional, built from realistic medical conversations and physician-written rubrics that grade accuracy, safety, and whether the model knows when to tell someone to seek care. In one comparison across 3,500 responses, OpenAI says physicians rated GPT-5.5 Instant above both older models and answers written by physicians working with internet access but no AI, with fewer missed red flags and referrals.
What it means for operators
Read the headline with one hand on the brake: these are OpenAI’s own evals, scored against OpenAI’s own rubrics. A vendor grading its own homework is a starting point, not proof.
The honest use is the boring one. A stronger free model is genuinely handy for understanding a lab result or drafting questions before an appointment, but it still gets things confidently wrong, and health is exactly where a wrong answer costs the most. Keep the verification habit: treat the output as a prompt for a real clinician, not a verdict.