The claim that AI has become strong at diagnosis while doctors remain better at choosing treatment is already out of date on its own evidence. The 78% figure cited in The Conversation reflects whether OpenAI’s o1-preview included the right diagnosis somewhere in a differential list of difficult New England Journal of Medicine cases—not whether it named the correct diagnosis first. And the same research found the model outscored physician groups on a benchmark explicitly designed to test treatment and follow-up planning. That distinction changes how Windows users should read the latest wave of health features in ChatGPT. AI is not licensed to make a diagnosis or prescribe a treatment, and it has not proven that it improves patient outcomes. But the neat dividing line between “AI for diagnosis” and “humans for management” no longer matches the published benchmark results.
The original analysis, written by University of Virginia physician and educator Andrew Parsons and distributed by The Conversation, is right about a central practical point: deciding between clinically acceptable options requires the patient’s own priorities, constraints and tolerance for risk. Its broader conclusion—that doctors are still better at weighing those options—goes further than the studies it invokes support.

Doctor consults a patient beside glowing AI-powered medical diagnostics and health data visualizations.The 78% o1 result is not a 78% diagnosis rate​

The headline number comes from a Science paper published on April 30, 2026, led by Peter Brodeur and colleagues at institutions including Beth Israel Deaconess Medical Center and Stanford. It evaluated OpenAI’s o1-preview reasoning model across several clinical-reasoning tasks and against historical physician benchmarks.
On 143 published NEJM clinicopathological conference cases, o1-preview included the correct answer in its differential diagnosis 78% of the time. Its top-one diagnosis was correct in 52% of cases. Those are both meaningful results, particularly for notoriously difficult cases, but they answer different questions.
A differential is the ranked list of possible explanations a clinician generates before deciding what evidence is needed next. Including a disease somewhere on that list can be useful, but it is not the same as confidently identifying the disease, deciding how urgent it is, ordering the right tests, ruling out dangerous alternatives or communicating uncertainty to a patient.
The paper also tested 76 real emergency-department cases in stages as more information became available. Its o1 model’s diagnoses were rated exact or very close in 67.1% of cases from the initial triage information, 72.4% after the emergency physician encounter, and 81.6% at hospital admission. Two attending physicians scored lower at each stage in that limited comparison.
Those results are impressive. They are also a long way from evidence that a consumer chatbot should become the first or final decision-maker for a sick child, chest discomfort, shortness of breath or any other potentially urgent symptom. The study was a retrospective, text-based evaluation of reasoning outputs, not a clinical trial in which AI independently treated patients and patient outcomes were measured.

The management-reasoning evidence contradicts the headline​

The material omission in The Conversation piece is that the Brodeur study did not stop at diagnosis. It included a management-reasoning experiment based on five “Gray Matters” cases, asking what clinicians should do next: which tests to order, which treatment path to choose, what to monitor and how to manage risk.
According to Stanford’s 2026 AI Index, which summarized the underlying study data, o1-preview reached a median management-reasoning score of 86%. GPT-4 alone scored 42%; physicians using GPT-4 scored 41%; and physicians using conventional resources scored 34%.
In other words, the study cited to establish AI’s diagnostic strength also found o1-preview substantially ahead of the human comparison groups on a structured test of clinical management. Calling treatment selection a remaining human advantage therefore cannot be presented as a conclusion from that research.
There is an important limitation: these were scored vignettes with expert-developed rubrics, not a years-long record of real patients facing cancer, heart failure, family obligations, insurance limits, adverse effects and changing goals. A rubric can assess whether a plan addresses the facts supplied to it. It cannot fully capture whether a patient later regrets surgery, declines radiation after seeing a friend suffer, loses access to transport, or changes their mind after a spouse becomes ill.
But that limitation applies to the physicians in the benchmark as well. It supports a narrower, defensible conclusion: we do not yet know whether AI-assisted management improves real-world patient outcomes. It does not establish that physicians currently outperform AI on management reasoning.
A 2025 randomized controlled trial in Nature Medicine reached an even more nuanced result. Ninety-two practicing physicians were assigned either GPT-4 plus normal resources or normal resources alone for five complex management cases. The AI-assisted physicians scored 6.5 percentage points higher than those without the model, but there was no statistically significant difference between the GPT-4-assisted physicians and GPT-4 by itself.
That is not evidence of a durable human lead. It is evidence that, in a simulated management exercise, the combination worked better than conventional reference tools, while the model alone performed comparably to the clinicians using it.

“It depends” is patient context, not an AI blind spot by definition​

Parsons’ prostate-cancer examples make an ethical and clinical argument rather than an experimental finding. Two 68-year-old men with similar slow-growing, early-stage tumors may rationally choose different paths: immediate treatment or active surveillance. One may be healthy but unable to live with uncertainty; another may have advanced heart failure and place greater weight on avoiding the burdens of treatment.
That is precisely how shared decision-making should work. It does not follow, however, that an AI system is incapable of structuring that decision or reflecting patient preferences once those preferences are provided.
Google DeepMind’s AMIE research, published in Nature in July 2026, addresses that gap directly. The company’s research system was tested in a blinded, virtual OSCE-style study involving 100 multivisit patient scenarios, 21 primary-care physicians and five specialties. The work was aimed at longitudinal management: tracking changes over time, responding to treatment, applying guidelines and handling medication decisions.
The researchers describe AMIE’s performance as physician-level, not superhuman, and the study remains a simulated assessment rather than deployed care. Still, it undercuts the claim that treatment planning is inherently beyond conversational AI. The field is moving toward systems designed to collect missing context, revisit plans over several encounters and explicitly account for patient preferences.
The more honest dividing line is operational. A chatbot can only reason over the information it receives, and it may misunderstand, invent details, overlook an urgent clue or provide a polished answer that sounds more certain than the evidence permits. A physician can examine a patient, detect distress, obtain tests, make referrals, respond to complications and accept legal and professional responsibility for the plan.
Those capabilities are not merely sentimental additions to medicine. They are how a management plan turns into care.

ChatGPT Health has arrived before outcome evidence has​

The timing matters because OpenAI began rolling out Health in ChatGPT to eligible U.S. users 18 and older on July 23, 2026. It is available on the web and iOS, meaning Windows users can access it through a browser even though the dedicated Health experience is not described as a Windows-native feature.
ChatGPT Health can connect, with permission, to medical records, Apple Health and supported health services. OpenAI says the feature can help users compare results over time, interpret visit summaries, prepare for appointments and understand patterns in their data. It also says connected health records and Health conversations are not used to train its foundation models or target advertising.
OpenAI is unusually explicit about the boundary: Health is intended to support, not replace, medical care and is not intended for diagnosis or treatment. Its newer ChatGPT for Clinicians product likewise says clinicians remain responsible for care decisions and should independently verify information.
That boundary is more than a disclaimer. The Food and Drug Administration’s final Clinical Decision Support Software guidance, issued in January 2026, draws distinctions between software that may fall outside device regulation and software functions that meet the definition of a medical device. Consumer-directed tools that influence diagnosis or treatment create different regulatory and safety questions from a clinician-facing reference system whose basis can be independently reviewed.
The FDA’s authorized AI-device list is dominated by bounded products, especially radiology and cardiovascular tools, with defined intended uses and premarket review. A general-purpose chatbot that reads a user’s narrative, records and wearable data does not fit neatly into that model. No public outcome trial establishes that directing patients through a broad conversational health system leads to safer care than ordinary access to clinicians and established triage services.

What a responsible AI health workflow looks like​

For consumers, the best current use of ChatGPT Health or a similar tool is preparatory rather than substitutive. It can turn scattered records into questions worth taking to an appointment, summarize a long discharge note, identify changes in a trend, explain unfamiliar terminology and help a caregiver keep track of instructions.
It should not be used as the sole gatekeeper for deciding whether troubling symptoms require urgent care. The problem is not simply that an AI may miss a diagnosis. It is that a delayed decision can cause harm even when the chatbot’s most likely diagnosis was reasonable.
The research supports using advanced models as a second set of eyes for clinicians and as an organizational tool for patients. It does not support treating a fluent answer as a care plan.
The real news is not that doctors have retained an uncontested treatment-planning advantage. The published record says the contest has already reached management reasoning, and recent models have performed at or above physician benchmarks in controlled tests. What remains unproven is more consequential: whether these systems, deployed with real patients and real responsibility, make the next decision safer.

References​

  1. Primary source: themercury.com
    Published: 2026-08-04T13:00:05+00:00
  2. Related coverage: news.med.virginia.edu
  3. Related coverage: washingtonpost.com
  4. Related coverage: thecrimson.com
  5. Related coverage: eweek.com
  6. Related coverage: news-medical.net
  7. Related coverage: public-pages-files-2025.frontiersin.org
  8. Related coverage: help.openai.com
  9. Related coverage: openai.com
  10. Related coverage: openai.com
  11. Related coverage: fda.gov
  12. Related coverage: help.openai.com