A June Nature Medicine study from NYU Langone Health found that GPT-5.2, Gemini 3.1 Pro and Claude Opus 4.6 outperformed OpenEvidence and UpToDate Expert AI across medical knowledge tests, clinician-alignment prompts and 100 real physician queries. The result matters for hospitals and clinicians adopting “medical-grade” chatbots: a specialized brand is not, by itself, proof of better answers or safer deployment.
As STAT reported on July 29, the paper has sparked unusually intense debate in health AI. That is understandable. Clinical tools have been marketed as a safer alternative to broad consumer and enterprise models, often on the strength of curated literature, retrieval systems and medical-specific workflows.
On 500 MedQA questions, Gemini led at 97.4% accuracy, followed by GPT at 94.2% and Claude at 90.2%. OpenEvidence scored 89.6%, while UpToDate Expert AI scored 88.4%.
The gap widened in HealthBench, where GPT scored 88.0 out of 100, compared with 62.6 for OpenEvidence and 61.3 for UpToDate. In the blinded evaluation of 100 de-identified, real-world clinician questions, the three frontier models formed a statistically higher tier than the two clinical products. UpToDate also declined to answer 19% of queries, more than the other tested systems.
That does not mean GPT, Gemini or Claude should be treated as autonomous clinicians. Crucially, the researchers found no statistically significant difference in harmful-content or hallucination flags among the tested systems. Better average benchmark performance is not a guarantee that an individual response is reliable enough to direct treatment.
NYU Langone’s authors point to retrieval-augmented generation, or RAG, as one possible source of weakness. Retrieval can ground a model in current evidence, but it can also distract it with irrelevant material or expose weaknesses in how the base model reconciles conflicting sources. Because OpenEvidence and UpToDate Expert AI are proprietary, the researchers could not inspect the underlying prompts, retrieval pipelines or model architectures.
The study also has real limitations. Clinical products were tested through browser interfaces rather than APIs, and the HealthBench portion used LLM-based judges, including frontier systems similar to the products under evaluation. The authors therefore describe their clinician-reviewed real-query test as the stronger evidence—and emphasize that this is a fast-moving snapshot, not a permanent ranking.
Any deployment that handles clinical questions should require identity controls, audit logs, data-loss safeguards, clear retention terms and a defined human-review workflow. A HIPAA-compatible instance, as used by NYU Langone for the source questions, is not the same thing as a green light to paste patient data into an unmanaged public chatbot.
The next purchasing conversation should focus on the exact task: literature lookup, note drafting, patient-message triage, differential diagnosis support or medication guidance. Each needs its own local evaluation, with representative cases, clinician review and a process for measuring errors after rollout.
The immediate lesson from the NYU Langone paper is not that general-purpose AI has won medicine. It is that doctors and health systems should trust measured performance and accountable workflows, not category labels.
As STAT reported on July 29, the paper has sparked unusually intense debate in health AI. That is understandable. Clinical tools have been marketed as a safer alternative to broad consumer and enterprise models, often on the strength of curated literature, retrieval systems and medical-specific workflows.
The benchmark result is clear, but not a clinical verdict
On 500 MedQA questions, Gemini led at 97.4% accuracy, followed by GPT at 94.2% and Claude at 90.2%. OpenEvidence scored 89.6%, while UpToDate Expert AI scored 88.4%.The gap widened in HealthBench, where GPT scored 88.0 out of 100, compared with 62.6 for OpenEvidence and 61.3 for UpToDate. In the blinded evaluation of 100 de-identified, real-world clinician questions, the three frontier models formed a statistically higher tier than the two clinical products. UpToDate also declined to answer 19% of queries, more than the other tested systems.
That does not mean GPT, Gemini or Claude should be treated as autonomous clinicians. Crucially, the researchers found no statistically significant difference in harmful-content or hallucination flags among the tested systems. Better average benchmark performance is not a guarantee that an individual response is reliable enough to direct treatment.
“Clinical AI” is a product category, not a safety certification
The study’s most useful finding may be its challenge to a convenient purchasing assumption. A tool can use medical retrieval, cite journals and sit behind a healthcare-focused interface while still producing an incomplete, unclear or occasionally unsafe synthesis.NYU Langone’s authors point to retrieval-augmented generation, or RAG, as one possible source of weakness. Retrieval can ground a model in current evidence, but it can also distract it with irrelevant material or expose weaknesses in how the base model reconciles conflicting sources. Because OpenEvidence and UpToDate Expert AI are proprietary, the researchers could not inspect the underlying prompts, retrieval pipelines or model architectures.
The study also has real limitations. Clinical products were tested through browser interfaces rather than APIs, and the HealthBench portion used LLM-based judges, including frontier systems similar to the products under evaluation. The authors therefore describe their clinician-reviewed real-query test as the stronger evidence—and emphasize that this is a fast-moving snapshot, not a permanent ranking.
Windows admins should recognize the governance problem
For IT teams, this is less about choosing a winning chatbot than it is about avoiding a familiar enterprise mistake: confusing a vendor’s vertical positioning with independently validated performance.Any deployment that handles clinical questions should require identity controls, audit logs, data-loss safeguards, clear retention terms and a defined human-review workflow. A HIPAA-compatible instance, as used by NYU Langone for the source questions, is not the same thing as a green light to paste patient data into an unmanaged public chatbot.
The next purchasing conversation should focus on the exact task: literature lookup, note drafting, patient-message triage, differential diagnosis support or medication guidance. Each needs its own local evaluation, with representative cases, clinician review and a process for measuring errors after rollout.
The immediate lesson from the NYU Langone paper is not that general-purpose AI has won medicine. It is that doctors and health systems should trust measured performance and accountable workflows, not category labels.
References
- Primary source: STAT
Published: 2026-07-29T08:30:00+00:00
Can doctors trust clinical AI? The complicated issue of LLM benchmarks | STAT
As researchers benchmark the safety and accuracy of clinical AI tools against ordinary chatbots, industry debates their findings.www.statnews.com