AI chatbots have become the newest front door to health information, offering instant explanations of symptoms, medications, test results, nutrition, and mental wellbeing concerns—but the growing habit of treating conversational AI as a medical adviser is creating risks that are easy to miss behind a confident, helpful-sounding answer.
A recent report from WPBF 25 News says that 17% of adults use AI chatbots for health-related questions at least monthly, citing survey findings discussed by Consumer Reports health reporter Kevin Loria. Tools including ChatGPT, Microsoft Copilot, and Google Gemini are appealing precisely because they can translate complex information into plain English, respond at any hour, and sustain a back-and-forth conversation without a waiting room, copay, or appointment.
Those strengths are real. They are also what make the problem more serious.
A chatbot can turn a confusing lab value into a digestible explanation, help a patient assemble questions for an upcoming appointment, or summarize a long medical leaflet. But it can also invent a dosage detail, mistake a dangerous symptom for a minor condition, overstate the likelihood of a diagnosis, or give incompatible answers to nearly identical prompts. And because the answer is fluent, structured, and personalized in tone, users may not recognize when the system has crossed from helpful explanation into unreliable medical guidance.
For Windows users, this issue is no longer confined to a browser tab. AI assistants are increasingly woven into search, office software, smartphones, operating systems, and productivity workflows. That proximity makes health questions feel like another ordinary prompt. They are not.
The reasons people turn to an AI chatbot for health information are not difficult to understand. Healthcare can be expensive, slow, intimidating, and full of dense terminology. A person who wakes up worried about a new symptom may get an immediate explanation from a chatbot in seconds, while a clinic appointment may be days or weeks away.
As Loria noted in the WPBF report, generative AI systems can be “quite good” at delivering substantial amounts of information in a form that is comprehensive and easy to understand. That is a meaningful advantage over a conventional search result page, which often leaves people to compare inconsistent sources and decipher medical jargon without context.
Examples include:
The difference matters because actual medical practice involves more than matching words in a prompt to patterns in a training dataset. Clinicians obtain a history, weigh what has and has not been said, perform examinations, interpret tests in context, identify uncertainty, and remain accountable for decisions. A general-purpose chatbot sees only the text it is given—and may misunderstand even that.
An answer may accurately explain a condition while missing a key exception. It may offer a reasonable list of possible causes but put the most alarming one first, escalating anxiety. Or it may offer reassurance without having enough information to safely rule out urgent possibilities.
Large language models are optimized to generate plausible language. They do not independently verify each sentence against a patient’s medical record, physical examination, current local care pathways, or an authoritative clinical database. When they lack knowledge, they can still produce an answer that sounds polished and complete.
The risk becomes worse when the user has no independent way to evaluate the response. Someone asking a chatbot to define a common medical term may be able to catch an odd statement. Someone asking whether a medication combination is safe, whether chest discomfort can wait, or whether a child’s symptom requires urgent care may not.
A major randomized study published in Nature Medicine found that when members of the public interacted with large language models for medical decision support, users struggled to communicate all necessary contextual information to the system. The models identified relevant conditions in a limited share of cases, and the study found that only about 34% of the possible conditions suggested during interactions were correct on average. The researchers also documented inconsistent handling of similar descriptions of potentially serious symptoms. Nature Medicine
That finding is especially important because real people do not present medical concerns in tidy exam-style prompts. They forget details, describe symptoms imprecisely, lead with an assumption, omit medications, or fail to recognize what is relevant. A model cannot reliably compensate for information it was never given.
But an AI system’s apparent confidence does not necessarily track its accuracy. The Nature Medicine study reported examples of users anthropomorphizing the systems and treating a confident tone as evidence of reliability. Nature Medicine
This is a familiar human-computer interaction problem with a new level of realism. A search engine returns links and encourages comparison. A chatbot delivers a single synthesized response in the voice of an informed assistant. That presentation can create an illusion that the hard work of evaluating evidence has already been done.
It has not.
Some studies find that generative AI can provide high-quality explanations in specific, controlled tasks. Others show promise when AI is used to support professionals working within structured systems. Yet those results do not establish that a general-purpose chatbot is safe for unsupervised health advice in everyday use.
A systematic review in JAMA Network Open examined studies evaluating large language models that provide health advice and found substantial variation in how those studies reported clinical accuracy. The review’s conclusion was not that the systems lack all value; rather, it highlighted that the evidence base is inconsistent enough to make sweeping claims about medical performance unreliable. JAMA Network Open
Patients describe symptoms ambiguously. They may change the story after follow-up questions. Multiple conditions can be present at the same time. Age, pregnancy status, medication history, allergies, test results, access to care, and local emergency numbers may all change what safe advice looks like.
Researchers at the University of Oxford emphasized this gap after a large user study found that public-facing LLM medical guidance could be inaccurate and inconsistent. The researchers argued that standardized testing alone cannot establish whether systems are safe for use by the public in high-stakes healthcare scenarios. University of Oxford
That is a crucial warning for technology enthusiasts who are accustomed to evaluating software by features, speed, and benchmark scores. In healthcare, a system that is right most of the time may still be unacceptable if its failures are unpredictable, difficult to recognize, or capable of changing someone’s treatment decision.
A Stanford Medicine-led study examining AI support for physicians found that clinician-AI collaboration has potential, but also exposed an important weakness: after doctors provided an assessment, the AI tended to agree with them even when instructed to reason independently. Stanford Medicine
That result illustrates two things at once. First, AI can contribute to clinical workflows where human judgment, patient records, and formal oversight are present. Second, even in a professional setting, systems can show automation bias and confirmation behavior rather than acting as a reliable independent check.
For consumers, the implication is straightforward. If trained clinicians still need to scrutinize AI output, ordinary users should be even more cautious about treating a chatbot’s response as a diagnosis, prescription, or triage decision.
People often enter highly sensitive details into AI chatbots: symptoms, photographs, medication lists, genetic risks, sexual-health questions, mental-health concerns, reproductive information, insurance details, lab results, and notes copied from a patient portal. They may do so because the chatbot feels private, immediate, and nonjudgmental.
That feeling can be misleading.
As the WPBF report notes, consumer AI services are not automatically subject to the same privacy rules that apply to a doctor’s office. In the United States, the familiar HIPAA framework covers specific kinds of healthcare organizations and their business associates; it does not automatically apply to every app, website, or consumer service that receives health-related information.
The Federal Trade Commission warns that health apps and services may use sensitive information for research, advertising, sharing with other companies, or even sale in some circumstances, and that such tools may not be covered by HIPAA in the same way a healthcare provider is. Federal Trade Commission
The U.S. Department of Health and Human Services explains that HIPAA protections apply to covered entities and business associates, not to all private companies handling health-adjacent data. If an organization does not meet those definitions, it is not required to comply with HIPAA’s rules. HHS
HHS further notes that health information voluntarily entered into an app that is not offered by or on behalf of a HIPAA-regulated entity is not protected by HIPAA merely because the information originated in a medical record. HHS
That does not mean every chatbot provider mishandles health data. It means users should not assume physician-level confidentiality from a consumer AI interface without reading the service’s current privacy terms, data controls, retention policies, and opt-out options.
That includes:
That approach preserves the chatbot’s practical strengths while recognizing the limits of its reliability, privacy protections, and awareness of personal context.
A practical verification routine is:
The danger is not only false reassurance. False alarm is also harmful. A chatbot can encourage unnecessary panic, push users toward a rare explanation, or reinforce a self-diagnosis that changes how they interpret every subsequent symptom.
A chatbot may provide general educational information, but it should not be the final authority on whether to begin, stop, split, combine, skip, or adjust medication. A pharmacist is often the most accessible qualified professional for medication-specific questions.
The Nature Medicine study’s finding of inconsistent responses to similar descriptions of serious symptoms demonstrates why chatbot triage is an unsafe foundation for decision-making. Nature Medicine
If a situation appears urgent, severe, rapidly worsening, or potentially life-threatening, people should contact local emergency services or a qualified medical professional rather than spend more time refining a chatbot prompt.
But health is also the category where the standard for error must be radically higher.
A wrong answer about a spreadsheet formula may waste ten minutes. A wrong answer about a symptom, drug interaction, or mental-health emergency can shape a decision with real consequences. The AI industry’s tendency to present assistants as broadly capable conversational partners increases the risk that users will not recognize when a routine query has become a high-stakes medical one.
The most promising future for AI in healthcare is therefore likely to be structured, accountable, and clinician-connected. That could include tools that summarize approved records for physicians, assist with documentation, help patients understand vetted information, support accessibility, or flag questions that need human review. It should not mean asking consumers to distinguish safe advice from dangerous fabrication on their own.
Yet the same systems can produce inaccurate answers with an authoritative tone, respond inconsistently to similar situations, and collect sensitive health information outside the protections people may expect. Research into public use of medical chatbots reinforces that concern: real-world interaction is messier than benchmarks, users may over-trust the technology, and even seemingly capable systems can fail when context is incomplete. University of Oxford
The safest rule is simple: use ChatGPT, Microsoft Copilot, Google Gemini, and similar tools to understand information and prepare better questions—but not to diagnose conditions, make treatment decisions, change medication, or replace professional medical care. In health, the final safeguard should be a qualified human who can see the full picture, test assumptions, protect confidentiality, and take responsibility for the advice given.
A recent report from WPBF 25 News says that 17% of adults use AI chatbots for health-related questions at least monthly, citing survey findings discussed by Consumer Reports health reporter Kevin Loria. Tools including ChatGPT, Microsoft Copilot, and Google Gemini are appealing precisely because they can translate complex information into plain English, respond at any hour, and sustain a back-and-forth conversation without a waiting room, copay, or appointment.
Those strengths are real. They are also what make the problem more serious.
A chatbot can turn a confusing lab value into a digestible explanation, help a patient assemble questions for an upcoming appointment, or summarize a long medical leaflet. But it can also invent a dosage detail, mistake a dangerous symptom for a minor condition, overstate the likelihood of a diagnosis, or give incompatible answers to nearly identical prompts. And because the answer is fluent, structured, and personalized in tone, users may not recognize when the system has crossed from helpful explanation into unreliable medical guidance.
For Windows users, this issue is no longer confined to a browser tab. AI assistants are increasingly woven into search, office software, smartphones, operating systems, and productivity workflows. That proximity makes health questions feel like another ordinary prompt. They are not.
The appeal of an always-available health assistant
The reasons people turn to an AI chatbot for health information are not difficult to understand. Healthcare can be expensive, slow, intimidating, and full of dense terminology. A person who wakes up worried about a new symptom may get an immediate explanation from a chatbot in seconds, while a clinic appointment may be days or weeks away.As Loria noted in the WPBF report, generative AI systems can be “quite good” at delivering substantial amounts of information in a form that is comprehensive and easy to understand. That is a meaningful advantage over a conventional search result page, which often leaves people to compare inconsistent sources and decipher medical jargon without context.
Where chatbots can be genuinely useful
Used with strict limits, AI can make health information easier to approach. The most defensible uses are largely organizational and educational rather than diagnostic.Examples include:
- Translating medical language into everyday English.
- Generating a checklist of questions to ask a doctor, pharmacist, nurse, or specialist.
- Explaining the difference between a screening test and a diagnostic test.
- Summarizing publicly available guidance from reputable health organizations.
- Helping a user prepare a concise timeline of symptoms, medications, allergies, and prior treatments for a clinical visit.
- Explaining what unfamiliar terms in an after-visit summary generally mean.
- Suggesting trusted sources to consult, such as government health agencies and major medical organizations.
The difference matters because actual medical practice involves more than matching words in a prompt to patterns in a training dataset. Clinicians obtain a history, weigh what has and has not been said, perform examinations, interpret tests in context, identify uncertainty, and remain accountable for decisions. A general-purpose chatbot sees only the text it is given—and may misunderstand even that.
Why fluent answers can be dangerously persuasive
The central problem is not that every chatbot answer is wrong. It is that a partially correct answer can be more dangerous than an obviously nonsensical one.An answer may accurately explain a condition while missing a key exception. It may offer a reasonable list of possible causes but put the most alarming one first, escalating anxiety. Or it may offer reassurance without having enough information to safely rule out urgent possibilities.
Large language models are optimized to generate plausible language. They do not independently verify each sentence against a patient’s medical record, physical examination, current local care pathways, or an authoritative clinical database. When they lack knowledge, they can still produce an answer that sounds polished and complete.
Hallucinations are not merely an academic problem
Loria’s warning in the WPBF report is blunt: AI systems can fabricate information, and those fabrications can be hard for users to detect. In health contexts, that can mean invented studies, nonexistent drug interactions, unsupported claims about treatments, or false assertions about a person’s symptoms.The risk becomes worse when the user has no independent way to evaluate the response. Someone asking a chatbot to define a common medical term may be able to catch an odd statement. Someone asking whether a medication combination is safe, whether chest discomfort can wait, or whether a child’s symptom requires urgent care may not.
A major randomized study published in Nature Medicine found that when members of the public interacted with large language models for medical decision support, users struggled to communicate all necessary contextual information to the system. The models identified relevant conditions in a limited share of cases, and the study found that only about 34% of the possible conditions suggested during interactions were correct on average. The researchers also documented inconsistent handling of similar descriptions of potentially serious symptoms. Nature Medicine
That finding is especially important because real people do not present medical concerns in tidy exam-style prompts. They forget details, describe symptoms imprecisely, lead with an assumption, omit medications, or fail to recognize what is relevant. A model cannot reliably compensate for information it was never given.
Confidence is not calibration
People often equate a detailed explanation with expertise. Chatbots amplify that instinct by using clear headings, empathetic language, caveats, and a conversational style that can feel more attentive than a rushed appointment.But an AI system’s apparent confidence does not necessarily track its accuracy. The Nature Medicine study reported examples of users anthropomorphizing the systems and treating a confident tone as evidence of reliability. Nature Medicine
This is a familiar human-computer interaction problem with a new level of realism. A search engine returns links and encourages comparison. A chatbot delivers a single synthesized response in the voice of an informed assistant. That presentation can create an illusion that the hard work of evaluating evidence has already been done.
It has not.
Research shows the evaluation problem is still unsettled
The evidence on AI chatbots in healthcare is neither uniformly negative nor ready to justify broad consumer reliance. That nuance matters.Some studies find that generative AI can provide high-quality explanations in specific, controlled tasks. Others show promise when AI is used to support professionals working within structured systems. Yet those results do not establish that a general-purpose chatbot is safe for unsupervised health advice in everyday use.
A systematic review in JAMA Network Open examined studies evaluating large language models that provide health advice and found substantial variation in how those studies reported clinical accuracy. The review’s conclusion was not that the systems lack all value; rather, it highlighted that the evidence base is inconsistent enough to make sweeping claims about medical performance unreliable. JAMA Network Open
Benchmarks do not equal bedside safety
A chatbot may perform well on standardized questions, medical licensing-style tests, or curated case descriptions. Real-world care is more complicated.Patients describe symptoms ambiguously. They may change the story after follow-up questions. Multiple conditions can be present at the same time. Age, pregnancy status, medication history, allergies, test results, access to care, and local emergency numbers may all change what safe advice looks like.
Researchers at the University of Oxford emphasized this gap after a large user study found that public-facing LLM medical guidance could be inaccurate and inconsistent. The researchers argued that standardized testing alone cannot establish whether systems are safe for use by the public in high-stakes healthcare scenarios. University of Oxford
That is a crucial warning for technology enthusiasts who are accustomed to evaluating software by features, speed, and benchmark scores. In healthcare, a system that is right most of the time may still be unacceptable if its failures are unpredictable, difficult to recognize, or capable of changing someone’s treatment decision.
AI may be more useful with clinicians than instead of them
There is a more constructive interpretation of the evidence: AI can be more valuable as a tool used within professional care than as a replacement for it.A Stanford Medicine-led study examining AI support for physicians found that clinician-AI collaboration has potential, but also exposed an important weakness: after doctors provided an assessment, the AI tended to agree with them even when instructed to reason independently. Stanford Medicine
That result illustrates two things at once. First, AI can contribute to clinical workflows where human judgment, patient records, and formal oversight are present. Second, even in a professional setting, systems can show automation bias and confirmation behavior rather than acting as a reliable independent check.
For consumers, the implication is straightforward. If trained clinicians still need to scrutinize AI output, ordinary users should be even more cautious about treating a chatbot’s response as a diagnosis, prescription, or triage decision.
Health privacy is the second major risk
Accuracy is only half of the story. The other major concern is privacy.People often enter highly sensitive details into AI chatbots: symptoms, photographs, medication lists, genetic risks, sexual-health questions, mental-health concerns, reproductive information, insurance details, lab results, and notes copied from a patient portal. They may do so because the chatbot feels private, immediate, and nonjudgmental.
That feeling can be misleading.
As the WPBF report notes, consumer AI services are not automatically subject to the same privacy rules that apply to a doctor’s office. In the United States, the familiar HIPAA framework covers specific kinds of healthcare organizations and their business associates; it does not automatically apply to every app, website, or consumer service that receives health-related information.
The Federal Trade Commission warns that health apps and services may use sensitive information for research, advertising, sharing with other companies, or even sale in some circumstances, and that such tools may not be covered by HIPAA in the same way a healthcare provider is. Federal Trade Commission
HIPAA is not a universal privacy label
A common misunderstanding is that any health-related information is automatically protected by HIPAA. That is not how the law works.The U.S. Department of Health and Human Services explains that HIPAA protections apply to covered entities and business associates, not to all private companies handling health-adjacent data. If an organization does not meet those definitions, it is not required to comply with HIPAA’s rules. HHS
HHS further notes that health information voluntarily entered into an app that is not offered by or on behalf of a HIPAA-regulated entity is not protected by HIPAA merely because the information originated in a medical record. HHS
That does not mean every chatbot provider mishandles health data. It means users should not assume physician-level confidentiality from a consumer AI interface without reading the service’s current privacy terms, data controls, retention policies, and opt-out options.
What not to paste into a chatbot
A prudent rule is to avoid entering details that could identify a person or expose sensitive records unless the service is explicitly approved by a healthcare provider or organization for that use.That includes:
- Full names, addresses, phone numbers, and dates of birth.
- Insurance member IDs and billing details.
- Medical record numbers and patient portal credentials.
- Complete lab reports, imaging reports, or discharge summaries containing identifiers.
- Photos that reveal identifying features, tattoos, documents, or location details.
- Information about another person who has not consented to its sharing.
- Highly sensitive reproductive, mental-health, substance-use, or genetic information.
A safer way to use ChatGPT, Copilot, or Gemini for health questions
The responsible model is not “never use AI.” It is use AI for preparation and comprehension, not diagnosis or treatment decisions.That approach preserves the chatbot’s practical strengths while recognizing the limits of its reliability, privacy protections, and awareness of personal context.
Use the chatbot as a question-building tool
Instead of asking, “Do I have condition X?” a safer prompt would be:Instead of asking, “Can I stop taking this medication?” a safer alternative is:“What questions should I ask my clinician about persistent fatigue, and what information should I bring to the appointment?”
Instead of pasting a complete report and asking for a diagnosis, a better use is:“Help me prepare questions for my pharmacist about side effects, interactions, and what to do if I miss a dose.”
These prompts steer the system toward education and communication, where errors are less likely to become immediate health actions.“Explain the general purpose of a lipid panel in plain language, using publicly available medical information.”
Verify before acting
For health information that could influence a decision, users should cross-check claims with reputable medical sources and, where appropriate, a qualified professional.A practical verification routine is:
- Ask for the source. Require the chatbot to identify a primary medical authority, guideline, or peer-reviewed research—not just make a claim.
- Open the original source. Do not accept a citation at face value. AI tools can generate inaccurate, irrelevant, or nonexistent references.
- Check the date and scope. Health guidance evolves, and recommendations may differ by age, country, pregnancy status, medical history, and medication.
- Compare at least two authoritative sources. Prefer government health agencies, recognized medical associations, established hospitals, and licensed clinicians.
- Bring unresolved questions to a clinician or pharmacist. This is particularly important for diagnoses, prescriptions, symptoms that are worsening, interactions, pregnancy, children, and mental-health crises.
The special danger of diagnosis, prescriptions, and triage
Certain categories of health questions should be treated as out of bounds for consumer chatbots as decision-makers.Diagnosis
A chatbot cannot examine a patient, check vital signs, order tests, observe appearance or movement, assess a rash under proper lighting, or recognize subtle changes that a trained clinician might detect. It can offer a list of possibilities, but lists are not diagnoses.The danger is not only false reassurance. False alarm is also harmful. A chatbot can encourage unnecessary panic, push users toward a rare explanation, or reinforce a self-diagnosis that changes how they interpret every subsequent symptom.
Medication and dosage changes
Drug questions are especially risky because a safe answer can depend on exact dosage, formulation, kidney or liver function, age, pregnancy status, other prescriptions, over-the-counter products, supplements, allergies, and timing.A chatbot may provide general educational information, but it should not be the final authority on whether to begin, stop, split, combine, skip, or adjust medication. A pharmacist is often the most accessible qualified professional for medication-specific questions.
Urgency and emergency decisions
Triage is among the most consequential uses of AI health advice because a mistaken recommendation can delay urgent care or send someone into unnecessary distress.The Nature Medicine study’s finding of inconsistent responses to similar descriptions of serious symptoms demonstrates why chatbot triage is an unsafe foundation for decision-making. Nature Medicine
If a situation appears urgent, severe, rapidly worsening, or potentially life-threatening, people should contact local emergency services or a qualified medical professional rather than spend more time refining a chatbot prompt.
What this means for Windows users and the AI ecosystem
Windows users have valid reasons to be enthusiastic about AI. Copilot and competing tools can make computing more approachable, summarize documents, brainstorm content, organize information, and reduce friction in daily tasks. Health literacy is one area where that accessibility can be genuinely beneficial.But health is also the category where the standard for error must be radically higher.
A wrong answer about a spreadsheet formula may waste ten minutes. A wrong answer about a symptom, drug interaction, or mental-health emergency can shape a decision with real consequences. The AI industry’s tendency to present assistants as broadly capable conversational partners increases the risk that users will not recognize when a routine query has become a high-stakes medical one.
The most promising future for AI in healthcare is therefore likely to be structured, accountable, and clinician-connected. That could include tools that summarize approved records for physicians, assist with documentation, help patients understand vetted information, support accessibility, or flag questions that need human review. It should not mean asking consumers to distinguish safe advice from dangerous fabrication on their own.
The bottom line: AI can inform, but it cannot take responsibility
AI chatbots are useful because they lower the barrier to asking questions. They can make medical language less intimidating, help people prepare for appointments, and point users toward information they might otherwise struggle to find.Yet the same systems can produce inaccurate answers with an authoritative tone, respond inconsistently to similar situations, and collect sensitive health information outside the protections people may expect. Research into public use of medical chatbots reinforces that concern: real-world interaction is messier than benchmarks, users may over-trust the technology, and even seemingly capable systems can fail when context is incomplete. University of Oxford
The safest rule is simple: use ChatGPT, Microsoft Copilot, Google Gemini, and similar tools to understand information and prepare better questions—but not to diagnose conditions, make treatment decisions, change medication, or replace professional medical care. In health, the final safeguard should be a qualified human who can see the full picture, test assumptions, protect confidentiality, and take responsibility for the advice given.
References
- Primary source: WPBF
Published: 2026-07-27T21:54:00+00:00
- Related coverage: ox.ac.uk
New study warns of risks in AI chatbots giving medical advice
The largest user study of large language models (LLMs) for assisting the general public in medical decisions has found that they present risks to people seeking medical advice due to their tendency to provide inaccurate and inconsistent information. The results have been published in Nature...www.ox.ac.uk
- Related coverage: communications.yale.edu
AI chatbots give inaccurate medical advice says Oxford Uni study
PDF documentcommunications.yale.edu