GPT-4 has scored higher on diagnostic reasoning tasks than doctors working with AI, but that result does not establish that a chatbot can deliver better medical care. The difference points to a harder problem: people and models can influence each other in ways that weaken reasoning, while real clinical decisions depend on complete records, communication and accountability.

What the diagnostic trials found

A randomized trial published in npj Digital Medicine tested diagnostic workflows with 70 practicing doctors. Both ways of incorporating AI improved accuracy over conventional medical resources. Yet GPT-4 working alone scored numerically higher than the doctors using it collaboratively. The reported results do not provide a standalone GPT-4 percentage for that comparison.

Study Participants Reported result
npj Digital Medicine diagnostic workflow trial 70 practicing doctors Accuracy was 85% with AI as a first opinion, 82% with AI as a second opinion and 75% with conventional resources. GPT-4 alone scored numerically higher than the collaborative workflows.
JAMA Network Open diagnostic reasoning trial 50 physicians A standalone large language model scored roughly 16 percentage points higher than the physician comparison group.
Nature Medicine study of AI assistants for the public 1,295 participants Models alone identified conditions in 94.9% of scenarios, while people given AI access identified illnesses less accurately than people without it.

A large language model, or LLM, generates responses from patterns learned during training. These studies measured performance on defined tasks, not whether an AI system could take responsibility for a patient’s care. They also show that adding a capable model to a person’s workflow does not automatically preserve the model’s standalone score.

Split scene of abstract AI reasoning and an anonymous clinician at a blank display, with diverging paths

▲ The gap between model and clinician reasoning

Why a strong model can become a weaker partner

One explanation is sycophancy: a conversational model may agree with an assumption built into a question, even when that assumption is wrong. A clinician or a member of the public may also anchor on an early possibility and give too little attention to alternatives. The model can then reinforce the initial idea rather than provide an independent check.

Workflow experience matters, too. Many clinicians in early trials had never used conversational AI chatbots. That makes the result a test of people using a tool in a particular workflow, not just a contest between medical knowledge held by humans and machines. It does not, by itself, show which factor caused the performance gap.

The Nature Medicine finding makes the distinction especially clear. High condition-identification accuracy from models operating alone did not carry over to people using those models. A reassuring or confident reply can shape a user’s next question, just as a leading question can shape the reply. Neither exchange substitutes for a clinical assessment.

Medical records create a different kind of risk

Even a model that handles a short diagnostic case well can struggle with a long patient history. In a JAMA Network Open study of Gemini 2.5 Pro hospital-course summaries, omission of important clinical details was the predominant reported error. Summarizing a lengthy record requires choosing what to leave out; a fluent summary can make those choices difficult to notice.

Long records also create context rot, a term for a model’s difficulty keeping track of relevant information across extensive material. Dates, sequences of events and changing laboratory values may be especially hard to retain accurately. Another problem is chart lore: an unverified assumption repeated across clinical notes until it appears to have a substantial history behind it. An AI system that reads those notes may repeat the assumption rather than independently establish whether it is true.

An anonymous clinician examines a long timeline of blank medical record sheets with gaps and repeated marks

▲ Omissions and repeated assumptions in records

These limitations matter because a missing detail or an inherited error can change the apparent meaning of a case. An AI-generated record summary is therefore a draft to check against the underlying information, not a replacement for the record or a clinician’s judgment.

Safety is broader than an accuracy score

A Stanford Medicine safety benchmark called noHARM assessed 20 LLMs across 4,249 clinical management options. It found that combining the outputs of distinct models produced safer, more robust advice than relying on a single model. Agreement across models may help expose inconsistencies, but it does not give any model professional responsibility for a decision.

The human side of the workflow needs protection as well. An observational study in The Lancet Gastroenterology & Hepatology found that endoscopists who had used AI assistance to detect polyps performed worse without it afterward than they had before using it. The finding raises a practical question for clinical teams: which skills should AI support, and which skills must clinicians continue to practice independently?

Medical competence also extends beyond retrieving a plausible answer. A trusted professional must explain uncertainty, communicate with a patient and remain accountable for decisions. Those obligations do not disappear when a model performs well on a reasoning exercise.

What to carry into a clinical conversation

The research points to a more limited, practical role for conversational AI: it can help explain the language in existing clinical notes or organize questions for an appointment. Sharing the original notes and objective lab results works better than a self-written summary, and symptoms should be described neutrally rather than framed as a suspected diagnosis, such as “Is this scabies?” Asking the model to assume a suspected condition is ruled out and to list alternative explanations can counter sycophancy. Running the same question through two or three separate models, such as ChatGPT, Claude and Gemini, can expose inconsistencies. Before a short appointment, a brief summary of the most important findings and questions can help prioritize the discussion. Its output should remain distinguishable from the original record, especially when it summarizes years of information. A polished explanation may still omit a finding or repeat an unsupported assumption.

The central lesson is not that doctors should ignore AI, or that a high-scoring model should make decisions alone. Standalone accuracy, effective collaboration and safe care are different measures. If AI helps prepare a discussion, bring the source information and unresolved questions to a qualified clinician. Diagnosis, treatment and other consequential medical decisions belong with a medical professional.