Mode
Text Size
Log in / Sign up

Public LLM web interfaces achieved 100% agreement with consensus risk categories in diabetes casesAI models match experts in identifying diabetes risk categories

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that LLMs achieved 100% agreement with consensus risk categories but require clinician supervision for safety.

This pilot benchmark evaluates the performance of five public LLM web interfaces in managing diabetes cases. The study utilized 20 translated, de-identified, and curated cases to assess how these models align with multidisciplinary consensus risk categories. The primary outcome showed 100% agreement (20/20 cases) with the frozen consensus risk category, with a Wilson 95% CI of 83.9%-100.0%.

Secondary outcomes included referral metrics and stability. The models identified a need for referral in 89/100 instances, correctly identified the referral specialty in 68/75 cases, and correctly identified the urgency of referral in 92/100 cases. Information-item and mandatory-measure coverage was reported at 100.0%. Stability across repeated generations was high for referral need (16/20), referral specialty (14/15), and urgency (19/20). No major safety errors were observed among 140 outputs.

Limitations include a 16.1-percentage-point interval below the observed ceiling for primary outcome precision and limited inter-reviewer reliability for completeness scoring. The authors emphasize that the absence of observed safety events in this small sample does not establish safety for routine clinical use. The findings suggest that while LLMs can provide consistent outputs, they require auditable, clinician-supervised decision support rather than autonomous deployment.

Managing diabetes requires accurate risk assessment to ensure patients get the right care at the right time. A new pilot study tested how well five different public AI web interfaces could handle these assessments. The researchers gave the AI 20 cases of diabetes data to see if the technology could match the decisions made by a team of medical experts.

The results were striking. The AI models achieved 100% agreement with the expert consensus on risk categories. They also performed well in identifying the need for referrals, the urgency of those referrals, and the specific type of specialist needed. In fact, the AI correctly identified the need for a referral in 89 out of 100 instances and correctly identified the urgency in 92 out of 100 cases.

While these results are promising, there are important limits to keep in mind. This was a small study of only 20 cases, and the researchers noted that a lack of safety errors in this small test does not mean the AI is safe for everyday use without a doctor. The study highlights that while AI can be a powerful tool for supporting decisions, it must be supervised by a clinician rather than used on its own.

What this means for you:
AI models matched expert risk assessments for diabetes, but they still need human supervision to be safe for patients.

Common questions

How accurate was the AI at identifying diabetes risk?

The AI models achieved 100% agreement with the expert consensus risk category across all 20 cases tested. This means the AI matched the experts' decisions perfectly in this specific study.

Can AI identify how quickly a patient needs a referral?

The AI performed well in identifying the urgency of a referral, correctly identifying the need for urgent care in 92 out of 100 cases. It also correctly identified the specific specialty needed for referral in 68 out of 75 cases.

Is it safe to use AI for diabetes management right now?

While no major safety errors were observed in the 140 outputs tested, the researchers noted that this does not mean the AI is safe for routine clinical use. They emphasize that AI should be used as a tool supervised by a doctor, not as an autonomous system.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedSep 2026
View Original Abstract ↓
BackgroundLarge language models (LLMs) may assist in the prevention of diabetes-related foot ulcers; however, their performance in classification may not translate effectively to context-dependent decisions.ObjectiveThis study aimed to evaluate the accuracy, clinical actionability, reproducibility, and safety of five public LLM web interfaces in the context of the International Working Group on the Diabetic Foot (IWGDF) 2023 risk stratification and preventive management.MethodsA prespecified, paired, blinded, noninterventional pilot benchmark was conducted using 20 translated, de-identified, curated cases structured in a fixed clinical sequence and evenly distributed across the four IWGDF risk categories. These input conditions represent an idealized, best-case benchmark rather than routine clinical documentation. Each interface assessed all cases under standardized no-search conditions. The primary outcome was exact agreement with a frozen multidisciplinary consensus risk category. Additional outcomes included screening frequency, need for referral, specialty and urgency of referrals, completeness, interreviewer reliability, repeated-generation stability, and safety.ResultsEach interface classified 20/20 cases in concordance with the frozen IWGDF reference (100%; Wilson 95% CI 83.9%-100.0%). The 16.1-percentage-point interval below the observed ceiling indicates limited precision and remains compatible with clinically meaningful error in new cases. Of 100 primary outputs, agreement was observed in 89/100 (89.0%) for referral need, 68/75 (90.7%) for referral specialty, and 92/100 (92.0%) for urgency. Median information-item and mandatory-measure coverage was 100.0% for both measures; however, completeness scoring had limited inter-reviewer reliability and should be interpreted cautiously. In an exploratory, hypothesis-generating four-case repeated-generation substudy, all-three-generation consensus concordance was observed in 16/20 interface-case combinations for referral need, 14/15 eligible combinations for specialty, and 19/20 combinations for urgency. These descriptive counts are not estimates of failure probability or tail behavior. Under the prespecified curated no-search benchmark conditions, no major safety errors were observed among 140 outputs; this absence of observed events does not establish safety in routine clinical use.ConclusionsPerformance was highest for structured guideline mapping, though reliability diminished in referral and individualized management across repeated generations. These findings highlight the necessity for auditable, clinician-supervised decision support instead of autonomous deployment.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.