Researchers evaluated how AI chatbots, including ChatGPT, Google Gemini, and Microsoft Copilot, provided advice on adjusting diabetes medications during the month of Ramadan. The study looked at 276 responses provided in both English and Arabic to check for accuracy, completeness, and safety.
The results showed that while 77% of the responses were fully consistent with medical guidelines, 11% of the responses were clinically harmful or contradictory. Notably, harmful responses were twice as common in Arabic than in English. The study also found that the accuracy and completeness of the advice did not vary significantly between the different chatbot platforms.
Because 30 out of 276 responses contained potentially harmful information, these tools are not reliable for making medical decisions. Consistency also fluctuated over time, making the advice unpredictable. These findings suggest that AI chatbots should only be used as a secondary resource and never as a replacement for professional medical advice from a doctor.