Mode
Text Size
Log in / Sign up

Dolphin bimodal LLM improves emotion recognition and patient satisfaction in education compared to text-only LLMTrial shows Dolphin AI improves patient education and satisfaction

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that bimodal LLMs like Dolphin may improve emotional alignment and patient satisfaction in education settings.

This randomized trial involved 555 patients across six departments and three centers to evaluate the impact of Dolphin, a bimodal large language model integrating text and audio cues, compared to a matched text-based large language model (LLM) in patient education.

Primary outcomes showed Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713, p < 0.001) and semantic consistency (84.9% vs. 82.1%, p < 0.001). Secondary outcomes indicated that Dolphin was associated with higher patient satisfaction (98.6% vs. 93.8%, p < 0.001), higher suggestion acceptance (76.1% vs. 58.9%, p < 0.001), and higher proactive disclosure (44.6% vs. 26.5%, p < 0.001). Additionally, Dolphin was associated with fewer 7-day unplanned recontacts (12.9% vs. 22.9%, p = 0.002).

No adverse events, serious adverse events, or unsafe recommendations were reported. While the randomized design suggests a causal link between the bimodal model and improved outcomes, the study evaluates a technical communication tool rather than a clinical treatment for a specific disease. Bimodal alignment may reduce communication failures in patient education settings.

Researchers conducted a randomized trial to compare two types of artificial intelligence in a medical setting. They compared Dolphin, a bimodal model that uses both text and audio cues, against a standard text-only model. The study included 555 patients across several medical departments and three different centers.

The results showed that the Dolphin model performed better in several areas. It was more accurate at recognizing emotions and remained more consistent in its meaning. Patients using the Dolphin system reported higher satisfaction levels and were more likely to accept suggestions. They also shared more information proactively compared to those using the text-only system.

Additionally, patients using the Dolphin system had fewer unplanned follow-up calls within seven days. No safety issues or harmful recommendations were reported during the study. While this trial shows that adding audio cues to AI can improve how patients receive information, it is important to note that the study tested a technical tool rather than a specific medical treatment.

What this means for you:
A bimodal AI model using text and audio cues showed better emotional alignment and higher patient satisfaction.

Common questions

How did the Dolphin AI compare to standard text-based AI?

The Dolphin model, which uses both text and audio cues, outperformed the text-only model in several areas. It showed higher accuracy in emotion recognition, better semantic consistency, and higher levels of patient satisfaction. It also led to more proactive disclosure from patients and fewer unplanned follow-up contacts within a seven-day period.

Was the use of the Dolphin AI safe for patients?

The study reported no safety concerns during the trial. No serious adverse events were identified, and the system did not provide any unsafe recommendations. The trial suggests that the bimodal approach is a safe way to enhance the way information is delivered to patients.

How did patient satisfaction change with the new AI?

Patients using the Dolphin system reported higher satisfaction levels compared to those using the text-only model. The study also found that patients were more likely to accept suggestions and share information proactively when using the bimodal system, which may help reduce communication failures.

Study Details

Study typeRct
Sample sizen = 555
EvidenceLevel 2
PublishedSep 2026
View Original Abstract ↓
BACKGROUND: Effective patient education requires accurate communication aligned with patients' emotional and semantical needs. Text-based large language models (LLMs) lack access to non-verbal cues, which may contribute to misaligned responses. METHODS: We evaluated emotional and semantic misalignment in a text-based LLM using 64,200 utterances from 16,583 patient education cases across six departments and three centers. Dolphin was developed integrating text and audio cues and evaluated through emotion recognition, semantic consistency assessment, branch-level ablations, and a double-blinded randomized trial against a matched text-based LLM comparator (Chinese Clinical Trial Registry: (ChiCTR2500095933). FINDINGS: The text-based LLM showed emotional misalignment in 36.7% of responses and semantic misalignment in 28.3% of cases, with higher misalignment under greater burden. Dolphin outperformed the text-based LLM in emotion recognition accuracy (0.886 vs. 0.713) and semantic consistency (84.9% vs. 82.1%; both adjusted p < 0.001). Ablations supported contribution of audio branches. Dolphin received higher expert ratings than the text-based LLM and human educators (all p < 0.001). In 555 patients, Dolphin was associated with greater patient satisfaction (98.6% vs. 93.8%), suggestion acceptance (76.1% vs. 58.9%; p < 0.001), proactive disclosure (44.6% vs. 26.5%; p < 0.001), and fewer 7-day unplanned recontact (12.9% vs. 22.9%; p = 0.002). No unsafe recommendations or safety events were identified. CONCLUSIONS: Compared with text-based LLM, Dolphin improved emotional-semantic alignment and patient-education outcomes, supporting bimodal alignment as a strategy for reducing misalignment-driven communication failures. FUNDING: National Natural Science Foundation of China, State Key Laboratory Special Fund, and Chinese Academy of Medical Sciences Innovation Fund.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.