Home›Ophthalmology› GPT-5 leads LLMs on cataract patient education quality but readability remains suboptimal
GPT-5 leads LLMs on cataract patient education quality but readability remains suboptimalAI models provide varying quality of information for cataract patients
Frontiers in MedicinePublished September 25, 2026DOI ↗Editorial oversight: Dr. Lars van Dijk, PhD · Surgical, Procedural & Diagnostic
AI-generated summary of the cited source, checked by automated accuracy review.
How we work
Share
Key Takeaway
Consider expert review before using LLM-generated cataract education materials.
This systematic review assessed the readability, information quality, and educational suitability of responses from five mainstream large language models (Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5) to 20 frequently asked questions about cataract. The review did not report a study population, sample size, or setting.
Significant differences were found among the models across the assessed outcomes (p < 0.05). GPT-5 achieved the highest scores on the c-PEMAT-P and GQS instruments, indicating superior information quality and educational suitability, but also had higher readability difficulty indices, as did Tongyi Qianwen. Content category influenced readability: responses on postoperative management and risk/prevention had better educational suitability, while those on surgical diagnosis, treatment, and preoperative management were more difficult to read. Correlations between quality indicators and readability-related indicators were generally weak.
The authors note limitations: LLM output quality varies by model and topic, and GPT-5 readability is not always at an ideal level. They emphasize that LLMs are not a replacement for expert review and that readability scores do not directly correlate with information quality.
For clinical practice, LLM-generated cataract patient education materials should undergo expert review, readability optimization, and patient-centered validation before being used in clinical settings.
How this fits prior evidence
This review extends prior coverage of cataract management by shifting focus from surgical techniques and IOL calculations to patient education tools. Previous items addressed capsular bag fixation in pediatric cataract surgery, new IOL formulas for post-RK patients, and combined phacoemulsification for glaucoma, all clinician-directed interventions. This review evaluates LLM-generated educational content, a distinct domain. It confirms that readability and quality are not interchangeable, echoing the caution in prior coverage that emerging tools require validation before clinical adoption.
When patients look for information about cataract surgery, they might turn to AI tools like GPT-5 or other large language models. However, a new review shows that these AI systems do not provide the same quality of information. The study found significant differences in how well different models explained the condition and the surgery.
While some models performed better in terms of information quality and educational suitability, others were harder to read. For example, topics like postoperative management were easier for patients to understand, while surgical diagnosis and treatment were often more difficult to follow. Even the top-performing models did not always hit the ideal mark for readability.
Because these tools vary so much in their accuracy and clarity, they cannot replace a doctor's advice. The study highlights that any information generated by AI for cataract patients must be reviewed by a medical expert and adjusted to ensure it is clear and helpful before it is used in a clinical setting.
What this means for you:
AI models vary in quality and clarity when explaining cataracts, requiring expert review before use.
Common questions
Are AI models reliable for cataract information?
AI models show significant differences in the quality and readability of their information. While some models performed better than others in educational suitability, the quality of the content varies depending on the specific model used. Because of these inconsistencies, AI-generated materials must be reviewed by a medical expert before being used for patients.
Which topics are hardest for AI to explain clearly?
The study found that surgical diagnosis, treatment, and preoperative management were more difficult to read. In contrast, information regarding postoperative management and risk prevention had better educational suitability. This means some parts of the cataract journey are harder for AI to explain simply than others.
Can I rely on AI instead of a doctor for cataract advice?
No, AI models are not a replacement for expert review in clinical settings. Because AI output quality varies by model and topic, and because readability scores do not always mean the information is high quality, you should always consult your doctor for medical advice regarding your vision.
BackgroundCataract is one of the main causes of visual impairment and reversible blindness worldwide, and it mainly affects the elderly population. Clinically accurate and sufficiently readable patient education materials play a crucial role in this regard. With the rapid development of large language models (LLMs), patients are increasingly obtaining health information generated by artificial intelligence (AI); however, the reliability of online medical information is often questionable. This study systematically evaluated the readability, quality, and educational suitability of mainstream LLMs when answering questions related to cataract.MethodsFive mainstream LLMs - Doubao, DeepSeek, Wenxin Yiyan, Tongyi Qianwen, and GPT-5 - were evaluated based on their responses to 20 frequently asked questions (FAQs) for cataract patient, which covered five subject categories. Text readability was assessed through multiple indicators, including the Coleman-Liau Index (CL), Linsear Write (LW), Automated Readability Index (ARI), Simple Measure of Gobbledygook (SMOG), Gunning Fog Index (GFOG), Flesch Reading Ease Score (FRES), and Flesch–Kincaid Grade Level (FKGL). Information quality and educational suitability were evaluated using the Global Quality Score (GQS) and the Chinese version of the Patient Education Material Readability Assessment Tool (c-PEMAT-P). Differences between groups were compared using one-way ANOVA and Kruskal-Wallis tests, with correlation analyses exploring relationships among indicators.ResultsThere were significant differences among LLMs in terms of readability, information quality, and educational suitability (all p < 0.05). GPT-5 had the highest c-PEMAT-P and GQS scores, however, several readability difficulty indices of GPT-5 and Tongyi Qianwen were also higher. There were significant differences in readability among different content categories. The postoperative management and risk/prevention topics tended to have better educational suitability, whereas surgical diagnosis, treatment, and preoperative management were more difficult to read. Correlation analysis demonstrated that the correlation between the quality indicators and the readability-related indicators is generally weak.ConclusionIn terms of generating educational materials for cataract patient, LLMs have potential, but the output quality varies depending on the model and the topic. GPT-5 performs best in terms of overall quality and educational suitability, but its readability is not always at an ideal level. Model selection has a crucial impact on the quality and educational suitability of the information, while the content topic mainly affects the language complexity. Therefore, for the responsible use of LLMs in cataract patient education, the cataract education materials generated by LLMs need to undergo expert review, readability optimization, and patient-centered validation before clinical application.