Debates about the medical use of artificial intelligence are often framed incorrectly. Critics mostly ask whether AI can already replace a doctor entirely, then treat every shortcoming as evidence that the technology cannot yet be taken seriously clinically. This is a misleading benchmark. The real question is not whether artificial intelligence is already an independent doctor, but how far AI models have progressed in certain aspects of medical reasoning.
Everyday, publicly available AI models are already remarkably strong in some areas of clinical work. A recent study published in JAMA Network Open examined 21 widely available large language models on clinical reasoning tasks, and their raw diagnostic performance was particularly strong. “Raw accuracy” means the proportion of responses judged correct in the clinical tasks concerned. In many cases, these systems can recognise patterns of disease, connect symptoms with likely conditions, interpret laboratory results, suggest treatment directions, organise medical text and formulate responses whose structure and content already resemble competent clinical reasoning.
The models performed particularly well once sufficient clinical information was available. Their accuracy on final diagnosis ranged from 61 to 91 percent, while they remained markedly weaker at the beginning of the diagnostic process, when they had to consider a broad range of possible conditions. This difference shows that AI is already strong at drawing a final conclusion from known information, but is not yet reliable enough to independently carry out the entire clinical reasoning process.
Since the release of GPT‑6 Astra, newer healthcare performance data have also become available. Astra scored 63.4 percent on HealthBench Professional, which is based on real clinical tasks, compared with 60.5 percent for GPT‑5.6 Sol, while its general hallucination rate fell from 12.2 percent to 4.2 percent. The new model is therefore more accurate and reliable, handles long documents and complex workflows more effectively, but these results still do not demonstrate that it can replace a physician as an independent clinical decision-maker.
Moreover, the JAMA study did not test AI systems developed specifically for hospital use, but publicly available “base” models such as ChatGPT. The models did not have full electronic health record integration. They had no dedicated clinical workflow behind them. They were not given hospital-level access to specialist database queries. They were not fine-tuned for medical specialties.
Consider an everyday example. We receive laboratory results containing values such as haemoglobin, CRP, blood glucose, creatinine, liver enzymes, cholesterol, TSH or vitamin D: many patients do not know what to do with this information. They see numbers, asterisks and reference ranges but do not understand what matters.
Today, however, they can enter these data into ChatGPT, or even photograph them, and ask: “Explain these laboratory results in plain language. Which values are outside the normal range? What might they indicate? What should I ask my doctor?”
A good AI model does not thereby become a doctor, and it should not provide a definitive diagnosis, but it can immediately do several useful things. It translates medical abbreviations into everyday language and highlights high or low values. It groups related findings: for example, inflammatory markers, kidney function, liver function, thyroid values, parameters suggesting anaemia or metabolic risks. It explains that a single abnormal value often does not mean disease in itself, while certain patterns may warrant medical attention. It can also provide a structured list of questions so the patient can return to their doctor better prepared.
For example, if CRP is elevated, the model can explain that this may indicate inflammation or infection but is not specific on its own. If fasting blood glucose and HbA1c are high, it may suggest discussing insulin resistance or diabetes with the doctor. If creatinine is high or eGFR is low, it can indicate that kidney function should be checked. If liver enzymes are elevated, it can list possible causes - such as medication effects, alcohol use, fatty liver, viral or metabolic problems - while making clear that interpretation depends on symptoms, medical history, medications and follow-up tests.
This is valuable in itself. Not because the chatbot becomes a doctor, but because it turns an opaque medical document into an understandable, organised initial interpretation. It helps patients ask better questions, better understand their condition and have a more meaningful consultation with their doctor. That alone is a significant practical breakthrough.
The limitations are real, however. These models remain weaker at the beginning of the diagnostic process, when information is scarce, uncertainty is high and several possible diseases need to be considered in parallel. They may tend to settle too early on one likely answer rather than maintaining a broad differential diagnosis. They need more advanced uncertainty management, stronger safety boundaries, better clinical context and closer integration with validated medical knowledge bases.
But this is a development challenge, not a dead end: future systems are expected to become stronger through structured questioning, retrieval of current medical guidelines, clinical calculators, combined analysis of visual and textual data, patient record integration and workflows operating under medical supervision.
In some well-defined medical tasks, AI may already match or exceed part of the performance of an average professional. The next genuine breakthrough, however, is likely to come not from another model alone, but from safely connecting AI with patient data, current guidelines, clinical calculators and physician oversight.
References
Rao AS et al.: Large Language Model Performance and Clinical Reasoning Tasks, JAMA Network Open, 2026.
Mass General Brigham: AI Remains Lacking in Clinical Reasoning Abilities, According to Study of 21 Large Language Models, 2026.
Peter Attia: Can AI models reason like clinicians?, 2026.
OpenAI: GPT‑6 Astra: A new generation of intelligence, 2026.
OpenAI: HealthBench Professional: Evaluating Large Language Models on Real Clinician Chats, 2026.
The models examined
The cited study examined 21 widely available large language models, including models from OpenAI, Anthropic, Google DeepMind, xAI and DeepSeek. Examples mentioned in public summaries included models such as GPT-5, Claude 4.5 Opus, Gemini 3.0 Flash/Pro, Grok 4 and DeepSeek models.
GPT‑6 Astra was not included in the JAMA study because the research was completed before the model was released. Astra’s results therefore come from separate benchmarks and are not directly comparable with the study’s PrIME-LLM scores.




