Skip to content
H
Unlocking infinite medical capacity.

Meet Harrison.Rad 1.5, our next-gen foundation model. Learn more →

Domain Beats Generalist Where It Counts: The Case for Specialised Models in Medical Imaging
Evidence

Domain Beats Generalist Where It Counts: The Case for Specialised Models in Medical Imaging

By Dr Jarrel Seah, Neuroradiologist and Chief Medical & AI Officer, Harrison.ai

A recent Nature Medicine paper reported that general-purpose large language models outperformed two clinical AI tools across several medical benchmarks. The convenient reading is that generalist models beat specialised medical AI, but I believe this to be too broad, particularly in radiology. 

The paper is also contested. OpenEvidence has raised concerns about benchmark contamination, scoring, and the description of the benchmark dataset. I am not going to adjudicate that dispute here. Even taking the result at face value, the more useful question is what kind of clinical AI was being evaluated, and how that evaluation was conducted. 

That question matters because clinical AI is no longer one thing. We now evaluate it with multiple-choice questions, simulated consultations, board-style examinations, reader studies, prospective trials, and arena-style comparisons. None of these reflect the totality of clinical practice.  

What the study tested 

The Nature Medicine study tested tasks that were text-based: licensing-style questions, clinician preference alignment, and de-identified clinical queries judged by reviewers. On those tasks, frontier models did better, which is meaningful. It says something about medical text knowledge and medical question-answering. But it does not tell you how these models read a scan. 

OpenEvidence and UpToDate Expert AI are clinical tools built around language models, retrieval, trusted medical content, and clinician-facing interfaces. This can save clinicians valuable time, help cite sources, and make medical knowledge easier to use. But these “specialised” models fundamentally perform the same type of task that today’s frontier LLMs are trained upon – the parsing of natural language queries, memorisation of facts, the use of tools to retrieve knowledge, and reasoning in the language domain.  

Why radiology is different 

Radiology is different because the information that matters is not conveyed in text, but in images. The clinical notes may reveal an extensive smoking history, but the cancerous lung nodule itself can only be found in the pixels. A guideline may tell you what to do with a finding, but it does not contain the visual signal you need to detect it. Frontier models can reason their way into a treatment plan, but reasoning cannot surface a finding it never saw in the image. 

To confidently detect these abnormalities requires training on hospital data, PACS data, radiologist reports and addendums, with the variation of real scanners and protocols. A frontier model can be extraordinarily capable and still have seen very little of the data that matters most for reading a scan. That is why a model trained on radiological imaging behaves so differently on an image-interpretation exam. 

Does the evaluation measure what it claims to? 

The deeper issue is one of construct validity, a problem every health system already applies to its own quality programs. Does the test accurately measure the specific abstract concept that you care about?  

One of the distinct features of radiology is that it has an unusual alignment between what can be tested in an examination and what the specialist actually does in practice. What makes a good physician is not just their ability to answer clinical questions, but also their ability to elicit information from patients, interact with other specialties, and deal with uncertainty. While assessments like HealthBench close the gap between real world practice and licensing examinations, they still present simulated patients as opposed to actual live human responses. On the other hand, a large part of what makes a good radiologist is perceptual expertise, diagnostic reasoning, where the artefacts that form a real clinical case can be reproduced exactly and sampled at scale.  

A text-based medical benchmark tests knowledge and reasoning in language. A radiology short case is closer to the real work of a radiologist: look at the images, find what is relevant, ignore what is not. Neither is a perfect stand-in for clinical practice, but they measure very different things, and only one of them tells you about image interpretation. Radiology fellowship examinations are still just examinations, not clinical practice itself, but they are far more task-matched to image interpretation than text-only medical questions. 

Testing Harrison.Rad 1.5 on radiology examinations 

We have tested Harrison.Rad 1.5 internally using mock examinations inspired by the Fellowship of the Royal College of Radiologists 2B Short Case examinations – a component of the UK radiology fellowship examination requiring image interpretation rather than text recall alone. 

The evaluation compared Harrison.Rad 1.5 with leading general-purpose models. Harrison.Rad 1.5 is a research-only foundation model, and Harrison.ai is seeking regulatory clearance in various markets for medical devices powered by Harrison.Rad 1.5. 

Harrison.Rad 1.5 achieved a median score of 86.5 against a pass threshold of 73.2. The strongest general-purpose model evaluated, GPT-5.4, scored 44. The other general-purpose models scored below 38. 

That is a large gap. It points to something health-system leaders should understand: intelligence is multidimensional, and applying the right benchmarks and tests is critical to picking the best model for the task. It is also consistent with the idea that a model trained on radiological imaging behaves differently from models trained primarily for language and broad multimodal use. 

It is still an internal result: we have not released the examination set, and Harrison.Rad 1.5 has not yet faced external scrutiny, so it should be treated as a useful stress test, not the final word. That scrutiny is exactly what we welcome external researchers to partner with us on. 

Independent evaluation of radiology AI 

Independent evaluations are beginning to test this in more relevant settings. In a 2026 American Journal of Roentgenology study, radiologists at Stanford and Mass General Brigham compared four AI systems on 212 chest radiographs. Harrison.Rad 1 produced the reports radiologists most often accepted, up to 75.5% versus 16 to 57% for the others, and it was rated highest for quality and preferred most often. Separately, at the 2025 American College of Radiology Annual Meeting, a challenge run with Mass General Brigham had 113 radiologists blind-rate Harrison.Rad 1 reports acceptable 65.4% of the time, against 79.6% for radiologist-written reports. These evaluations settle nothing on their own, but they test something closer to the clinical task than any medical text benchmarks, and Harrison.Rad 1.5 will be held to the same standard. 

Specialisation only helps if it gives the model something real: access to different data, competence in a different modality, or both. For medical text, that advantage may be small. For radiology image interpretation, it is much larger. 

Why the gap is unlikely to close 

General-purpose models will improve in all domains, including medicine. The question is what ingredients they need in order to improve. Radiology performance depends on clinical imaging data linked to reports, corrections, follow-up, outcomes, pathology, scanner variation, protocol variation, and difficult cases encountered in real clinical use. That data is generated inside health systems. It is not sitting in a public web crawl. 

Our models do not learn from deployment in real time. Like other regulated medical software, each version is trained on curated datasets, validated, locked, and released. Learning happens across development cycles: performance gaps are identified, data are curated, labels are improved, edge cases are added, and a new version is trained and validated. 

The Nature Medicine paper may be a useful result for text-based tasks, medical question-answering and sitting examinations, but the practice of radiology is a different task, and if you are choosing AI to read scans, you should insist it be evaluated as one. 

Dr Jarrel Seah, Neuroradiologist and Chief Medical & AI Officer, Harrison.ai

 

Harrison.Rad 1.5 and Harrison.Rad 1 are research-only foundation models, not medical devices regulatory-cleared for clinical use. Harrison.ai is seeking regulatory clearance for medical devices powered by these models in various markets. 

Meet Harrison.Rad 1.5

Experience what’s possible when foundation models are purpose-built for radiology.