Mode
Text Size
Log in / Sign up

LLM-based AI tools show large positive effect on empathy with SMD 1.02 in communicationAI tools show large gains in patient empathy scores

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that LLM-based tools significantly improve empathy scores and patient comprehension in communication tasks.

This meta-analysis synthesized data from various study designs to evaluate the impact of Large Language Model (LLM) based AI tools, such as GPT-3 and GPT-4, on physician-patient communication. The included studies involved a diverse population including patients with chronic conditions, healthcare professionals, and laypersons, totaling 2604 participants in the pooled analysis. The primary objective was to determine if LLM-generated responses could improve empathy compared to standard physician-generated content.

The interventions consisted of LLM-based AI tools and chatbots used to generate patient communications, while the comparators included human-written responses from physicians. In several specific implementations, GPT-4 was utilized to simplify complex medical information, such as pathology reports, for patient consumption. The study evaluated multiple dimensions of communication, including empathy, clarity, information quality, satisfaction, trust perceptions, and consultation time.

The primary outcome of empathy showed a large positive effect when LLM assistance was used, with a reported standardized mean difference (SMD) of 1.02 (95% CI 0.44-1.60). In direct comparisons between chatbot responses and physician responses, LLMs were rated significantly higher in empathy in 5 out of 6 instances. Specifically, one comparison showed a 45.1% success rate for chatbots versus 4.6% for physicians (OR ~9.8, P<0.001). In a specific study comparing ChatGPT-4 to human responses on a 5-point scale, ChatGPT-4 scored higher with a mean of 4.18 compared to 2.70 (P<0.001). Another neurology-focused study reported higher empathy scores for ChatGPT answers by +1.38 on the Consultation and Relational Empathy Scale (P<0.01).

Secondary outcomes highlighted significant improvements in patient understanding and efficiency. When GPT-4 was used to simplify pathology reports, patients showed increased comprehension scores of 7.98 compared to 5.23/10 for standard reports. This intervention also resulted in a reduction of 70% in consultation time (P<0.001). These results suggest that LLMs can effectively translate complex medical jargon into accessible information while potentially streamlining the clinical workflow.

Regarding safety and tolerability, specific adverse events or discontinuation rates were not reported in the included studies. However, the analysis noted that AI-generated replies were sometimes less concise or less readable for patients with low literacy levels. Furthermore, the study did not provide data on long-term trust impacts or the potential for inaccurate advice in certain contexts. These results compare favorably to traditional methods of patient education where physician time constraints often limit the depth of empathetic communication. While these findings are promising, they must be weighed against the fact that LLM tools are not a replacement for physician oversight. Methodological limitations include the lack of long-term trust assessments and the variety of study designs included in the meta-analysis.

Clinically, these results suggest that LLM-based chatbots can enhance communication by producing more empathetic and understandable responses. They may be particularly useful in simplifying complex reports or managing high-volume information. However, clinicians must remain aware that AI tools can generate overly lengthy content. Questions remain regarding the long-term impact on patient trust and the accuracy of AI-generated advice in specialized clinical scenarios.

When people live with chronic health conditions, the way a doctor communicates can make a huge difference. Patients often feel overwhelmed by complex medical terms or may feel that their concerns are not being fully heard. This research looks at how using large language models, which are types of AI like ChatGPT, can change the way doctors and patients talk to one another. The goal is to see if these tools can help make medical information easier to understand while making the interaction feel more supportive.

The researchers conducted a meta-analysis, which means they combined data from several different studies involving over 2,600 participants. These included people with chronic conditions, healthcare professionals, and members of the general public. They compared responses generated by AI tools against those written by human doctors to see how patients perceived the information in terms of empathy, clarity, and overall usefulness.

The results showed that when AI was used to help draft responses, there was a large positive effect on empathy scores. In several direct comparisons, the AI-generated answers were rated significantly higher for empathy than those written by humans. For example, one specific model scored much higher on scales measuring how well it addressed the patient's feelings and needs. Additionally, when AI was used to simplify complex medical reports, patients reported better understanding of their condition, and the time spent during consultations was reduced by about 70 percent.

While these results are promising, there are important things to keep in mind. Some studies found that AI responses could sometimes be too long or less readable for people who have difficulty with complex texts. Also, because this was a meta-analysis of various study types, we do not yet know how these tools affect long-term trust between patients and their doctors over many years. The researchers also noted that more large-scale, real-world trials are needed to see how this works in daily practice.

For patients today, this means that AI is not intended to replace your doctor. Instead, it can be a tool that helps your healthcare provider give you clearer information and more empathetic support. While the technology is still being tested for its long-term effects on the doctor-patient relationship, it shows potential for making medical communication more helpful and easier to navigate during difficult times.

What this means for you:
AI tools can help doctors provide more empathetic and clear information, but they do not replace human oversight.

Study Details

Study typeMeta analysis
Sample sizen = 2,604
EvidenceLevel 1
PublishedJul 2026
View Original Abstract ↓
BACKGROUND: Recent advances in large language models (LLMs) such as GPT-3/4 have spurred the development of artificial intelligence (AI) chatbots and advisory tools in medicine. These systems are posited to assist or augment physician-patient communication, potentially improving empathy, clarity, and responsiveness. However, their actual impact on communication outcomes remains uncertain. OBJECTIVE: This study aimed to systematically review and meta-analyze peer-reviewed studies (2020-2025) evaluating how LLM-based interventions affect physician-patient communication, including empathy, clarity, trust, and patient understanding. METHODS: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines, we searched PubMed/MEDLINE, Embase, Scopus, and Web of Science for studies published from 2020 to 2025 examining LLM or chatbot applications in clinical communication contexts. Eligible designs included randomized, observational, cross-sectional, and qualitative studies. Two reviewers (WHP and SR) independently screened titles or abstracts, assessed full texts, and extracted data on study design, population, LLM type, communication measures, and outcomes. We conducted a qualitative synthesis and random-effects meta-analysis, reporting pooled standardized mean differences or odds ratios with 95% CIs. RESULTS: From 312 records, 10 studies were included, all quantitative and predominantly cross-sectional. Populations ranged from patients with chronic conditions to health care professionals and laypersons. Outcomes assessed included empathy (8 studies), clarity or information quality (6 studies), satisfaction or usefulness (4 studies), and trust perceptions (2 studies). In 6 direct comparisons of AI- versus physician-generated responses, LLMs were rated significantly higher in empathy in 5 studies. One large study found that chatbot replies were judged empathetic in 45.1% of cases versus 4.6% for physician replies (odds ratio approximately 9.8, P<.001). Similarly, ChatGPT-4 answers scored higher in empathy on a 5-point scale than human-written responses (mean 4.18 vs 2.70, P<.001). One neurology study showed higher empathy scores (Consultation and Relational Empathy Scale +1.38, P<.01) for ChatGPT answers. Only 1 study found no significant empathy difference. LLM content was also longer and more information-rich, improving patient-perceived clarity and understanding. On the other hand, GPT-4 simplified pathology reports, increasing patient comprehension scores (7.98 vs 5.23/10, P<.001) and reducing consultation time by 70%. However, AI replies were sometimes less concise or less readable for low-literacy patients. In pooled analyses (k=4 studies; total evaluations N=2604), LLM assistance showed a large positive effect on empathy (standardized mean difference 1.02, 95% CI 0.44-1.60; random-effects model). Patient satisfaction results were mixed. No study directly assessed long-term trust. CONCLUSIONS: Current evidence suggests that LLM-based chatbots can enhance physician-patient communication by producing more empathetic, detailed, and understandable responses. These improvements may positively influence patient experience and engagement. However, LLMs may also generate overly lengthy or occasionally inaccurate advice, emphasizing the need for physician oversight. While meta-analytic findings are promising, robust randomized controlled trials, real-world and longitudinal studies are needed to confirm benefits, assess trust outcomes, and define optimal clinical integration strategies.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.