Mode
Text Size
Log in / Sign up

DeepSeek-R1 outperforms ChatGPT-4o in Chinese DDH caregiver guidance accuracyComparing Large Language Models for Providing Information on Hip Joint Conditions

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Consider LLM accuracy differences when recommending AI tools for DDH caregiver education.

This comparative study, aligned with guideline principles, assessed the performance of two large language models (LLMs), ChatGPT-4o and DeepSeek-R1, in answering 53 Chinese-language questions about developmental dysplasia of the hip (DDH). The questions were caregiver-oriented, and the primary outcome was clinical accuracy and guideline concordance. Secondary outcomes included educational usability, inter-rater reliability, and readability.

DeepSeek-R1 demonstrated higher clinical accuracy than ChatGPT-4o for caregiver-oriented questions, with mean scores of 4.70 (SD 0.15) versus 3.69 (SD 0.28), a difference that was statistically significant (P < 0.001). The study did not report effect sizes or confidence intervals, and no adverse events or tolerability data were collected.

The study's scope is limited to evaluating LLM outputs, not clinical outcomes in patients. It does not provide evidence on whether these models improve caregiver understanding or patient management. The authors did not report limitations, funding sources, or conflicts of interest.

For clinicians, this study suggests that LLM responses to caregiver questions about DDH may vary in accuracy, with DeepSeek-R1 showing an advantage in this Chinese-language context. However, given the lack of patient-level outcomes and the absence of reported limitations, these findings should be interpreted cautiously. LLMs should not replace professional medical advice, and further validation is needed before considering their use in patient education.

How this fits prior evidence

This comparative study extends prior coverage on AI applications in DDH by evaluating LLMs for caregiver-facing information. It complements the May 2026 meta-analysis on AI-assisted hip ultrasound, which focused on diagnostic accuracy, whereas this study addresses patient education. The finding that DeepSeek-R1 outperforms ChatGPT-4o in Chinese-language DDH questions adds a new dimension, but it does not directly confirm or contrast with prior findings on surgical planning or AVN risk factors. It addresses a gap in evaluating AI tools for caregiver communication, but the lack of patient outcomes limits its clinical applicability.

Doctors and researchers looked at how two different AI systems, ChatGPT-4o and DeepSeek-R1, answered questions about a hip condition called Developmental Dysplasia of the Hip. This condition affects how a baby's hip joint develops and can lead to problems later in life.

The study tested 53 specific questions that parents and caregivers might ask. The goal was to see which AI gave better medical advice that matched current health guidelines. The researchers checked for accuracy, how easy the information was to read, and how helpful it was for families.

The results showed that the DeepSeek-R1 model provided more accurate clinical information than ChatGPT-4o. While both tools can provide information, one was found to be more reliable when giving advice to caregivers. This helps doctors understand which tools might be safer for patients to use.

It is important to remember that these tools are not doctors. They are being tested to see how well they can explain medical facts. Parents should always talk to a real doctor before making any medical decisions for their children.

What this means for you:
DeepSeek-R1 provided more accurate medical information for hip conditions than ChatGPT-4o.

Common questions

Which AI chatbot was more accurate for hip dysplasia questions?

DeepSeek-R1 scored higher than ChatGPT-4o. On a 5-point scale, DeepSeek-R1 averaged 4.70, while ChatGPT-4o averaged 3.69. The difference was statistically significant, meaning it was unlikely to be due to chance.

What is developmental dysplasia of the hip?

It's a condition where a baby's hip joint doesn't form properly. The ball part of the joint may be loose or out of place. Early treatment can help, so getting accurate information is important for caregivers.

Can I use AI chatbots for medical advice about my child's hip?

This study suggests some AI models give better answers than others, but it didn't test real patient outcomes. AI should not replace professional medical advice. Always talk to your child's doctor for guidance on treatment and care.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedSep 2026
View Original Abstract ↓
BackgroundLarge language models (LLMs) are increasingly used for caregiver-facing health information, but their reliability in Chinese-language pediatric orthopaedics remains uncertain. This study evaluated whether responses to developmental dysplasia of the hip (DDH) questions were clinically accurate, aligned with Chinese guidance, and educationally usable.MethodsWe compared ChatGPT-4o and DeepSeek-R1 using 53 Chinese-language DDH questions, including 31 caregiver-oriented frequently asked questions and 22 guideline-derived items. Each model generated one response per question using a standardized single-turn, five-sentence prompt. Six blinded pediatric orthopaedic surgeons rated clinical accuracy and guideline concordance. Paired model comparisons, inter-rater reliability, and exploratory formula-based readability were assessed.ResultsDeepSeek-R1 had higher clinical accuracy than ChatGPT-4o across 31 caregiver-oriented questions (mean question-level score 4.70 [SD 0.15] vs. 3.69 [0.28]; P 
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.