top of page
Search

AI in Medicine Is Already Better Than Critics Admit


The debate about artificial intelligence in medicine is often framed in the wrong way. Critics ask whether AI can fully replace physicians today, then treat every limitation as proof that the technology is not clinically serious. That is a weak standard. The real question is not whether AI is already an autonomous doctor. The real question is how far ordinary, publicly available AI models have already moved into medical reasoning — and whether the latest studies are even measuring the strongest systems available.

The real story is not that AI failed medicine. The real story is that ordinary, publicly available AI models are already astonishingly good at parts of clinical medicine — and the study did not even test the newest and strongest models.


A recent JAMA Network Open study evaluated 21 off-the-shelf large language models on clinical reasoning tasks. The negative reading is obvious: these models still struggle with early-stage differential diagnosis, especially when patient information is incomplete. That weakness matters. In real medicine, missing a dangerous alternative diagnosis can be catastrophic.

But the positive signal is far more important than the headline suggests.

The models achieved remarkably strong raw diagnostic performance. Raw accuracy means the percentage of answers judged correct across clinical tasks. In practical terms, this is not abstract “AI hype.” It means these systems can often recognize disease patterns, connect symptoms with likely conditions, interpret laboratory findings, propose treatment directions, organize medical text, and produce answers that resemble competent medical reasoning. That is a massive leap from where consumer AI was only a few years ago.

The study found that the models were especially strong when enough clinical information was available. Final diagnosis and management were the best-performing areas. According to Mass General Brigham’s summary, all tested LLMs reached a correct final diagnosis more than 90% of the time when given all relevant case information. That is not a marginal result. It shows that general-purpose AI systems already have serious medical pattern-recognition capacity.

This point needs emphasis: these were not specialized hospital AI systems. They were general, off-the-shelf models. No full electronic health record integration. No dedicated clinical workflow. No hospital-grade retrieval system. No specialty-specific fine-tuning. In plain terms, a normal consumer chatbot can already produce surprisingly useful preliminary medical reasoning when given structured information. It must not replace a doctor, but dismissing it as trivial is intellectually sloppy.



A practical example makes this clear. Imagine an ordinary person receives a blood test result. Traditionally, they may see values such as hemoglobin, CRP, glucose, creatinine, liver enzymes, cholesterol, TSH, or vitamin D, without really understanding what they mean. Today, that person can copy the results into ChatGPT and ask: “Please explain these lab results in plain language. Which values are outside the normal range? What could they suggest? What should I ask my doctor?”

A competent AI model will not replace the physician, and it should not pretend to make a final diagnosis. But it can translate medical abbreviations into normal language, identify high or low values, group related findings, and explain possible medical significance. It may distinguish inflammation markers, kidney function, liver function, thyroid function, anemia markers, and metabolic risk. It can also generate a structured list of questions for the doctor.

For example, if CRP is elevated, the model may explain that this can indicate inflammation or infection, but is non-specific. If fasting glucose and HbA1c are high, it may suggest discussing diabetes or insulin resistance with a physician. If creatinine is elevated or eGFR is low, it may flag kidney function as an issue to review. If liver enzymes are high, it may suggest possible liver, medication, alcohol, metabolic, or viral causes — while clearly stating that interpretation depends on symptoms, history, medication use, and repeat testing.


This is already valuable. Not because the chatbot becomes a doctor, but because it turns an opaque medical document into an understandable, organized first interpretation. It helps the patient prepare for a better medical consultation. That alone is a major practical breakthrough.

The limitation is also clear. These models are still weaker at the beginning of the diagnostic process, where information is incomplete and uncertainty is high. They may converge too quickly on one likely answer instead of maintaining a broad differential diagnosis. They still need better uncertainty handling, better safety guardrails, better clinical context, and stronger integration with validated medical knowledge.

But this is a development problem, not a dead end. Future systems will likely improve through structured prompts, retrieval from current guidelines, clinical calculators, multimodal inputs, patient-record integration, and physician-supervised workflows.

The sober conclusion is: even ordinary AI models are already strong at final diagnosis and management, while still needing supervision for differential diagnosis and real-world clinical uncertainty. The direction of travel is unmistakable.


References

Rao AS et al., Large Language Model Performance and Clinical Reasoning Tasks, JAMA Network Open, 2026.Mass General Brigham, AI Remains Lacking in Clinical Reasoning Abilities, According to Study of 21 Large Language Models, 2026.Peter Attia, Can AI models reason like clinicians?, 2026.


Models tested

The reported study evaluated 21 off-the-shelf LLMs, including models from OpenAI, Anthropic, Google DeepMind, xAI, and DeepSeek; examples named in public summaries include GPT-5, Claude 4.5 Opus, Gemini 3.0 Flash/Pro, Grok 4, and DeepSeek models. GPT-5.5 was not included.

 
 
 

Comments


bottom of page