Imagine a busy clinic in a low-resource area, where a doctor needs help reading an X-ray but lacks the powerful computers and huge datasets that most medical AI relies on. A new review of current research shows that while AI models that combine language and vision are getting better at tasks like answering questions about images or generating reports, they often don't work well in these real-world settings.
The review looked at how these models are built and found that the most common approach, called cross-modal transformers, is very good at matching images with text, but it needs a lot of computing power and massive, carefully organized datasets. That's a problem for clinics with limited hardware and spotty internet.
The good news? The review points to smarter ways to make these models more practical. Techniques like parameter-efficient fine-tuning, which adjusts a model without retraining everything, and federated learning, which trains across multiple devices without sharing raw data, could help. These methods are more realistic for low-resource settings.
But here's the honest catch: this is a review of trends, not a clinical trial. It doesn't give specific numbers on how well any one model performs in a real clinic. The authors also note that there are still big gaps, like not enough data in many languages and models that aren't always well-calibrated. So while the path forward is clearer, we're not there yet.