Mode
Text Size
Log in / Sign up

AI chatbots provide clinically harmful or contradictory antidiabetic advice in 11% of responses during RamadanAI Chatbots Provide Inconsistent Advice for Managing Diabetes During Ramadan

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that 11% of AI chatbot responses for Ramadan medication adjustments were clinically harmful or contradictory.

This guideline evaluates the performance of three AI chatbots (ChatGPT, Google Gemini, and Microsoft Copilot) regarding antidiabetic medication adjustments during Ramadan. The evaluation focused on accuracy, completeness, safety, and reproducibility of responses provided in both English and Arabic.

Analysis of 276 responses revealed that 77% were fully consistent with guidelines, while 11% (30/276) were clinically harmful or contradictory. Harmful responses were approximately twice as common in Arabic (14%) compared to English (8%). While no significant differences were found between the specific chatbot platforms regarding accuracy or safety, the reproducibility of safety information over a two-week period was very low (kappa = 0.02, p = 0.81).

Limitations noted include the persistence of some harmful recommendations and fluctuating consistency over time. Given the risk of contradictory advice, these tools should not be used as a substitute for professional medical advice. They may only be utilized as a supplementary resource for clinicians or patients.

How this fits prior evidence

This guideline addresses a gap in the safety of digital health tools for diabetes management. While prior coverage has focused on pharmacological treatments like tirzepatide for specific syndromes and the use of sirolimus-eluting stents in diabetic patients, this evidence highlights the specific risks of using AI chatbots for medication adjustments during Ramadan. It confirms that AI-generated advice can contain harmful content, particularly in non-English languages.

Researchers evaluated how AI chatbots, including ChatGPT, Google Gemini, and Microsoft Copilot, provided advice on adjusting diabetes medications during the month of Ramadan. The study looked at 276 responses provided in both English and Arabic to check for accuracy, completeness, and safety.

The results showed that while 77% of the responses were fully consistent with medical guidelines, 11% of the responses were clinically harmful or contradictory. Notably, harmful responses were twice as common in Arabic than in English. The study also found that the accuracy and completeness of the advice did not vary significantly between the different chatbot platforms.

Because 30 out of 276 responses contained potentially harmful information, these tools are not reliable for making medical decisions. Consistency also fluctuated over time, making the advice unpredictable. These findings suggest that AI chatbots should only be used as a secondary resource and never as a replacement for professional medical advice from a doctor.

What this means for you:
AI chatbots can provide harmful medical advice for diabetes; they should not replace a doctor's guidance.

Common questions

Are AI chatbots safe for managing diabetes medication?

The study found that 11% of the responses, or 30 out of 276, were clinically harmful or contradictory. Because of these risks, AI chatbots should be used only as a supplementary resource and never as a substitute for professional medical advice from a healthcare provider.

Is there a difference in safety between English and Arabic responses?

The study found that harmful responses were about twice as common in Arabic as in English. Specifically, 14% of Arabic responses were harmful compared to 8% of English responses. You should always consult a doctor for medical guidance in any language.

Are some AI chatbots better than others for medical advice?

The study found no significant differences between ChatGPT, Google Gemini, and Microsoft Copilot regarding accuracy, completeness, or safety. All platforms showed some level of inconsistency, meaning none of the chatbots are currently reliable enough to replace a doctor's advice.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedAug 2026
View Original Abstract ↓
Patients with diabetes increasingly consult artificial intelligence (AI) chatbots for medical advice, including guidance on antidiabetic medication management during Ramadan fasting, because AI can simplify and summarize long, complex guidelines. Also, in hospital settings, these tools are being used in hospitals much faster than it takes to establish formal regulations and guidelines for their use. Evaluations of the accuracy, completeness, and reproducibility of such advice across languages are still lacking. Therefore, the study aims to evaluate and compare the accuracy, completeness, safety, and reproducibility of three widely used AI chatbots—ChatGPT, Google Gemini, and Microsoft Copilot—when providing antidiabetic medication adjustment advice during Ramadan in both English and Arabic. Twenty-three standardized clinical scenarios covering common antidiabetic regimens were presented to each chatbot in both English and Arabic. Each query was repeated to evaluate reproducibility, resulting in 276 responses scored. Responses were assessed against the International Diabetes Federation–Diabetes and Ramadan (IDF-DAR) Guidelines using a 0–2 accuracy scale, a 0–4 completeness scale, and a 0–3 safety scale. Overall, 77% of responses were fully consistent with the guideline, 12% were partially consistent, and 11% (30/276) contained clinically harmful or contradictory advice; harmful responses were about twice as common in Arabic as in English (14% vs. 8%). Completeness and safety were high, with medians at the observed ceiling. In the generalized linear mixed models, chatbots did not differ significantly in accuracy, completeness, or safety, and there was no significant main effect of language or chatbot × language interaction; the strongest signals were a chatbot effect on completeness (p = 0.068) and a language effect on safety (p = 0.064), both non-significant. Two-week reproducibility was fair for accuracy (weighted κ = 0.20, p = 0.009) and completeness (κ = 0.29, p = 0.001) and showed a very low κ in the safety scale (κ = 0.02, p = 0.81). AI chatbots demonstrated comparable performance in delivering guideline-based advice for diabetes management during Ramadan, with no significant differences in accuracy, completeness, or safety. While most responses aligned with the IDF-DAR guideline, some harmful recommendations persisted, and response consistency fluctuated over time. These results suggest that AI chatbots should serve as a supplementary resource rather than a substitute for professional medical advice.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.