Article: https://www.thelancet.com/journals/landig/article/PIIS2589-7500(26)00036-1/fulltext?rss=yes
Second article: https://www.thelancet.com/journals/landig/article/PIIS2589-7500(26)00030-0/fulltext?rss=yes
It sounds like the problem is the lack of standards around the models themselves, although that is more difficult to solve. My gut says that a single modality trained model will always out perform a generalist one. Excerpt here:
Artificial intelligence (AI) has rapidly transformed medical imaging, enabling automated detection, classification, and risk prediction across an increasing number of clinical applications. However, in clinical practice, physicians routinely integrate information across imaging modalities, specialties, and diagnostic contexts. In their Article published in The Lancet Digital Health, Yang Zhou and colleagues introduce MerMED-FM, a multimodal, multidisease medical imaging foundation model trained on more than 3·3 million images across seven imaging modalities and more than ten clinical specialties. The study represents an important step towards a generalist AI framework that is capable of supporting cross-disciplinary imaging interpretation.1
Most existing medical imaging AI models have been developed as task-specific tools, typically focusing on a single modality such as chest radiography, pathology slides, retinal photographs, or CT imaging.2,3 Although these specialist models can achieve excellent performance in narrow tasks, their fragmentation creates challenges for real-world use.4 Hospitals often need to maintain multiple AI systems across departments, each requiring separate validation, governance, and technical integration. Zhou and colleagues attempt to address this limitation by developing a unified vision-only foundation model that can interpret diverse medical images within a single architecture. By pretraining the model on large-scale multi-institutional datasets covering radiology, pathology, ophthalmology, dermatology, and ultrasound and combining self-supervised learning with a memory-based training mechanism, MerMED-FM demonstrates competitive performance across multiple clinical imaging benchmarks. The conceptual transition from fragmented specialist models to generalist medical foundation models is shown in the figure.
“My gut says that a single modality trained model will always out perform a generalist one.” i share that feeling
tho… it would require more digging.
ideas:- one extra model wont hurt. (unless it drains resources from other ones)
- is this for professional stuff?? or just for “wtvr” - OSS ?
- sharing/pooling dev effort can be efficient(cost saving)
I think all this AI technology is pretty much crap, but what I will say is that these tools are only useful when used in ensembles where multiple different models of disparate origins are consulted with the same question and then compared by a human to determine the quality of the results.
Not only are single modality trained models always going to outperform generalist ones given the way these tools work, they are also never going to be useful unless there are multiple discrete models trained to accomplish the same task being consulted.
What makes this all so broken is that nobody makes money by setting something up like Chatbot Arena where you can quickly feed the same data into multiple entirely different models, so the entire structure of the industry is problematic before even the first brick was laid down.
https://arena.ai/text/side-by-side
It is only when you input the same data into a tool like the above one that you can sketch a picture you can begin to trust. You can get quite good information out of very simple chatbots too because given comparisons against more complex ones it is far easier to establish trust, determine where innaccuracies and systematic issues lie, and route around problematic AI models if they are responding weirdly to a particular input of data.
I still don’t think it is worth the effort in most cases however unless large corporations buy up all the training data and destroy it/remove it from public access… at which point we will have a different more human problem.
i was digging into legal application(viability). there are many parallel concepts. its very interesting.
i am cautiously optimistic. (currently, is seems** that llm tech application, is hackity: “push the ‘on’ button… and we’re done!”)
- one extra model wont hurt. (unless it drains resources from other ones)



