AI can make a medical record easier to understand in minutes. It can also fail at a more urgent task: recognizing when someone needs emergency care. In a study published in Nature Medicine, ChatGPT Health did not recommend emergency care in 52% of standardized emergency scenarios. That contrast matters for anyone considering whether to put their health history into a chatbot: a useful summary is not the same thing as a safe clinical decision.
Where AI helps with health records
Electronic health records can be difficult to read as a whole. In one account, exporting a medical history from MyChart and uploading it to ChatGPT took about five minutes. The chatbot turned the records into a more comprehensible narrative and suggested topics to explore further. That ability may help a patient prepare questions, but a clinician still needs to assess what the information means for that person’s care.
AI is also taking on narrower work inside clinics. At a surgical practice at the Icahn School of Medicine at Mount Sinai, a specialized workflow checks which medications may need attention before an operation. The reported time savings were substantial:
| Weekly pre-operative review | Reported time |
|---|---|
| Staff medication cross-checking before the AI workflow | 5 to 10 hours |
| AI-assisted processing for the week’s surgical patients | About 5 minutes |
The workflow retains a required human review rather than allowing the system to make unchecked decisions. That distinction is essential: an AI tool can assemble information quickly, while a medical professional must verify it and make decisions about medication and surgery. Even with a review step, clinicians can grow accustomed to automated output and overlook something they would otherwise check.

▲ AI-assisted surgical preparation
These uses show the appeal of health AI without settling the question of trust. Organizing existing information is different from deciding how urgently a new symptom needs attention. The second task has consequences when a reassuring answer is wrong.
The triage test found dangerous gaps
The Nature Medicine evaluation tested ChatGPT Health with standardized medical vignettes—written descriptions of patient situations—and examined the level of care it recommended. In 52% of the emergency scenarios, it failed to recommend urgent emergency care.
One vignette described worsening breathing difficulty in a person with asthma who had used a rescue inhaler four times in 12 hours. The chatbot advised staying home for 24 to 48 hours. In another scenario involving suicidal thoughts, the addition of normal routine blood-test results caused the system to remove a crisis-support prompt and treat the situation as non-medical. Normal results in those tests did not resolve the danger described in the scenario.
OpenAI said the study evaluated an older model that higher-performing versions had replaced. It also said standardized vignettes did not reflect typical patient interactions. Subsequent testing across seven models and multi-turn conversations—exchanges with several back-and-forth messages—found persistent patterns of concern. Neither point makes the 52% figure a measure of every current chatbot or every real-world encounter. It does show why a fluent response should not be taken as proof of sound triage.
Good performance is not the whole of care
A Kaiser Family Foundation poll found that one in three U.S. adults, approximately 66 million people, had used AI for medical information. Patients also bring chatbot printouts to appointments. Those pages can lengthen a visit, but they can give a clinician a starting point for understanding what worries the patient.
The issue is not only whether AI can reach a correct answer. Medical care also involves acknowledging uncertainty, listening to someone in distress and taking responsibility for a decision. A system trained to give confident-sounding responses may obscure uncertainty just when a patient needs it explained. A physician can say that an answer is not yet known, discuss the available paths and remain accountable to the patient as circumstances change. AI may support that work; it does not provide the same human relationship.
That distinction becomes especially important when a diagnosis is life-altering or the choices ahead are uncertain. A patient may need more than a calculation of possible outcomes. They may need a medical professional to help interpret those possibilities in the context of their life.

▲ A patient considering chatbot guidance
A safer way to use a health chatbot
The FRESH framework offers a way to keep a chatbot in a supporting role. It is a guide to asking better questions, not a substitute for an examination or a diagnosis:
- Facts before framing: Describe symptoms and provide relevant test values before adding someone else’s reassuring opinion. Leading context can steer a chatbot away from the underlying concern.
- Reset your conversation: Start a new chat for a separate health issue so earlier exchanges do not shape the next answer.
- Examine the evidence and downsides: Ask what could make the answer wrong, what alternatives exist and what risks a doctor would consider.
- Seek care from a human: Use the response to prepare for a clinical conversation. A medical professional should make decisions about diagnosis, treatment and medications.
- High stakes? Trust yourself before a bot: Do not wait for a chatbot to validate an alarming symptom. Seek emergency evaluation for warning signs such as coughing blood, sudden numbness on one side, inability to walk or loss of bowel or bladder control.
Keep the clinician in the decision
AI’s speed makes it useful for reading records and helping clinical teams prepare, but speed and a plausible answer do not establish clinical safety. The triage findings put a clear limit on what patients should ask a chatbot to decide, even as newer models are released.
Use AI-generated summaries to organize questions and bring relevant information to an appointment. Ask a clinician to interpret findings and make care decisions. When symptoms may be urgent, bypass the chatbot and seek human medical evaluation instead.