Hacker News new | ask | show | jobs
by themantalope 33 days ago
Radiologist. I don’t read MR shoulder exams in my day to day practice, but from the few pictures shown , I can’t conclusively disagree with the original report.

These models are generally terrible at reading medical images. The amount of public training data on the internet compared to the number of scans a radiologist reads in training is minuscule. There’s obviously a ton of medical images in general but very few, and even fewer along with a report are available on the internet publicly for download.

There are vision language models coming out of research labs that are excellent in describing and localizing findings. Still at the level of a 1st or 2nd year radiology resident, but as we all say - this is the worst the models will ever be.

6 comments

Absolutely. It's very unfortunate that this post used the worst example possible of using LLMs for medical purposes.

General-purpose LLMs are _fantastic_ at medical diagnosis that do not involve imaging. I am completely convinced that given enough information and time, frontier models already outperform >90% of doctors on initial diagnosis of internal issues and suggesting medical tests to further reject or confirm the most likely theories. To the point where I'm eagerly waiting for the first hospital in the world that's willing to be open and honest about using them for that first step, and then proceeding from there. I'll be on a flight there as soon as one arrives.

At the same time, they're worse than useless at anything involving medical imaging. Asking them to interpret them is worse than trying to interpret them yourself as a layman. And you surely wouldn't interpret them yourself.

    > General-purpose LLMs are _fantastic_ at medical diagnosis that do not involve imaging.
Can you share the reasons that you believe this?

    > At the same time, they're worse than useless at anything involving medical imaging.
What is special about medical imaging that makes AI/LLMs specifically bad?
You can see it in just this PDF report.

It's multiple things. It never shows the subscapularis in the way that people actually look the tendon. It hyper fixates on the axial when I find the sagittal much more useful for subscapularis.

Figure 7. There's an arrow pointing "to the acromial undersurface". The arrow is not pointed to that location.

Figure 5. "thin bursal fluid". This is within physiologic variation, but is calling bursitis.

It keeps bringing up irrelevant normal things like the shape of the coracromipal arch, I assume because lots of websites have information about that as a patient focused possible cause for rotator cuff impingement.

I am reminded of the recent Stanford MIRAGE study which found that LLMs will happily hallucinate answers about medical images if the medical images are omitted.

https://arxiv.org/html/2603.21687v2

I don't understand why this is still confusing to people. The second "L" in LLM is language; these things are AWESOME at producing things that SOUND like language, including code. They have so much training data that it is almost always grammatically correct, and often makes sense. Extending this, it has obviously been trained on data containing phrases like "acromial undersurface" and "thin bursal fluid", and "coracromipal arch", in the context of shoulder injury and related imaging. BUT IT DOES NOT KNOW HOW TO DIAGNOSE ANYTHING. So, it SOUNDS like a radiologist or specialist, and might be in the ballpark of correct-ish-ness, but ultimately is a fancy Markov model.
lol yea I wasn’t going to put in a full dictation on the internet but clearly a lot of misinterpretations from the AI
> Can you share the reasons that you believe this?

Firstly, please keep in mind I'm talking about the entire doctor population of the world here. Not sure which particularly bubble of this earth you have experience with, but note how half the word's population lives in India/China/Indonesia/Pakistan/Nigeria/Brazil/Bangladesh/Russia. Now I do believe that it holds the same for e.g. Europe and non-China East-Asia, but still.

How many patients has the world-wide average doctor seen? How long have they been a doctor?

How many have they seen with the particular condition the patient has?

How much time do they spend listening to and reasoning about a patient? The median in the world is likely under 3 minutes.

How many real-world incentives do human doctors have to deal with?

Given infinite time and resources, and zero external incentives, maybe the median human doctor would outperform the LLM at this task. But this is completely detached from the real world.

> What is special about medical imaging that makes AI/LLMs specifically bad?

LLMs: Besides lack of training data as mentioned elsewhere, they're simply not trained for high-fidelity image processing in general. It's not limited to medical imaging. It's a bit like the "How many Rs in strawberry" thing, but worse.

As for "AI" in general, medical image analysis is a very active field. These tend to be purpose-built though, not general-purpose. It seems likely at some point they'll become mainstream, but there's still a way to go.

Not to mention shoulder MRs can be hard to read. Findings can be subtle and have to be interpreted in context with the exam, symptoms etc.
Yeah, medical computer vision is a (fascinating) field with a lot of ongoing research. SOTA models are highly specialized, and are only getting good enough to be used by actual doctors and patients. Using a general purpose LLM to do this is similar to giving a credit card to Openclaw and telling it to make you rich through the stock market & cryptos.
Somewhat. During residency I developed models for detecting liver tumors in MRI imaging. Like you said, highly specialized and a lot of manual work developing the dataset.

There are now open source open weight “foundation models” coming out of labs that are transformer or mamba based architectures. These will accelerate development.

I don't have insider information, but: if one of the AI companies really wants their models to become really good at this and publicly available datasets are scarce, they can probably just buy anonymized X-ray/MRI scans paired with the human doctor's diagnosis, and train on them. I don't know what the legal story is around this, but AI companies have near infinite money, so I'm sure they can buy their way around regulations (eg. by buying them from a less regulated country).
That’s a good question.

My understanding is that medical images are part of a patients record (so they must be available for the patient or other docs at the request of the patient) but whoever has collected the images does have some form of ownership. I’m not a lawyer and I have a cursory understanding of this. I believe it would be possible for an AI company to lease or get access to the data through a transaction but I suspect that it hasn’t happened (or happened publicly) due to fear of backlash.

For example I know that some companies like Tempus had access to imaging that corresponded to tumors which had been biopsies for sequencing and they were developing models in house.

Anecdotally, I've had Claude (Sonnet and Opus latest) consistently misread numbers from screenshots of my macro tracking app. Makes me skeptical of claims about its usefulness for anything requiring accurate image interpretation, let alone MRI analysis.
I can see how your thesis is valid.

Like OP, I also had a shoulder MRI, and asked two AIs for opinion (awaiting a follow up appointment to discuss the results).

They both insinuated much more serious problem than it was (as judged by an orthopaedic doctor).

No trolling here: Do you feel threatened by the advance of AI/LLMs with respect to your field? I would. I am a computer programmer, and it absolutely feels threatening.
I mostly do interventional radiology, which is more similar to a surgical specialty than a diagnostic radiology practice. We do a lot of procedures , see patients in our clinic, have a service that rounds on and follows up on patients etc. I don’t think AI will affect my job much in the next 10-20 years. Advancements in AI and robotics could potentially offload some of our work, but most of what’s out there is underwhelming, but we all know that can change fast.

Even for diagnostic, I’m not totally convinced AI is going to cause a lot of issues. From a medical-legal perspective I doubt AI companies are willing to take on the risk of misdiagnosis . When that happens, human rads will start to get phased out.

In the interim, I think the next 10 years could be a golden period for diagnostic rads. Rads will still be the ones doing the work and signing reports, but people who learn how to use the right combination of tools will become very productive and can make a great living. Eventually payers will re-align but early adopters who figure it out will have a leg up.

As a programmer, I don’t feel threatened by the technology itself, but I do feel threatened by the second-degree effects such as what the technology does to our field, especially in the wrong hands.