|
|
|
|
|
by yread
29 days ago
|
|
Another problem is that general models' performance just sucks. From an upcoming conf. talk (in pathology) where they ran 2 Medgemma models on 100 slides with known diagnosis: > Results: Full concordance with the reference diagnosis was 8% (27B) and 5% (1.5 4B; McNemar p=0.68), while partial matches were 29% vs 20% respectively (McNemar p=0.053). When correct diagnoses anywhere in the differential were counted, 51% (27B) vs 30% (1.5 4B), with 27B significantly superior (McNemar χ²=12.1, p=0.0005). Site-level performance varied widely (30–100%). Both models reported HIGH confidence in ~99% of cases irrespective of correctness. i.e. highly confident, wrong 95% of time. in 49% of cases the real diagnosis wasn't even on models' differential. Doctor can hardly improve using something they can safely assume to be just noise. https://ecp2026.abstractserver.com/programme/#/scientific/de... |
|
Not sure how that research compares to the claims being made by many that a second opinion via ai in the end led to changes in treatment. Likely people spent quite some time searching and figuring out. That would be a different and n=1 result. Don't have enough knowledge of that research to determine how much result can be gained when the models are managed in a way that produces better results.
And of course how much time/effort/cost that would take. How much is custom and how much is an automated programmable flow.