Mode
Text Size
Log in / Sign up

GPT-4o mini answers oculoplastic questions with mixed accuracy and readabilityAI Model Shows Mixed Results for Oculoplastic Patient Questions

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
GPT-4o mini showed limited clinical accuracy and completeness for oculoplastic patient questions.

A cross-sectional evaluation examined how GPT-4o mini responded to 33 patient questions spanning eight oculoplastic domains, including floppy eyelid syndrome, cicatricial ectropion, entropion, blepharoplasty, facial aging, brow lift, ptosis, and facial nerve function. Six raters scored each response across accuracy, completeness, readability, transparency, objectivity, and bias, yielding 198 total ratings.

Clinical accuracy was rated favorably in 51% of ratings (101/198), while completeness reached 42% (84/198). Transparency and objectivity fared better: 78% (155/198) and 76% (150/198) of ratings indicated agreement or strong agreement, respectively.

Bias was uncommon. Only 8% of ratings (16/198) indicated bias was present, whereas 65% (129/198) disagreed or strongly disagreed that bias appeared in the responses.

The authors note that findings apply specifically to GPT-4o mini and should not be generalized to other large language models. No funding or conflicts of interest were reported. GPT-4o mini showed limitations in clinical accuracy, completeness, and readability for oculoplastic patient questions, suggesting that current outputs may not yet meet clinical or readability requirements without expert review.

Researchers evaluated how well the GPT-4o mini AI model answered 33 questions regarding various eye conditions, such as floppy eyelid syndrome, ptosis, and facial aging. The study looked at how accurate, clear, and unbiased the AI's responses were when providing information to patients.

The results showed that the AI was only rated favorably for clinical accuracy in 51% of the cases. It was even less consistent with completeness, receiving favorable ratings in only 42% of the responses. While the AI was seen as objective and transparent by many users, its ability to provide complete and accurate medical information was limited.

Because the AI struggled with accuracy and completeness, these results suggest that patients should be cautious when using AI for medical information. These findings apply specifically to the GPT-4o mini model and may not apply to other AI systems. Always consult a medical professional for reliable information regarding eye conditions.

What this means for you:
The GPT-4o mini model showed limited accuracy and completeness when answering questions about eye conditions.

Common questions

How accurate was the AI at answering eye surgery questions?

The AI was rated favorably for clinical accuracy in 51% of the ratings. This means that in nearly half of the cases, the responses were not considered accurate enough for clinical use. Because of these limitations, patients should consult a doctor for reliable medical information.

Was the AI's information complete?

The AI was rated favorably for completeness in only 42% of the ratings. This indicates that the AI often failed to provide a full range of information regarding conditions like ptosis or floppy eyelid syndrome. You should always verify medical details with a healthcare professional.

Was the AI biased when answering questions?

The AI was generally seen as unbiased. In 65% of the ratings, participants disagreed or strongly disagreed that bias was present. However, the model still showed significant limitations in accuracy and completeness for many clinical topics.

Study Details

Study typeSystematic review
EvidenceLevel 1
PublishedSep 2026
View Original Abstract ↓
PurposeThis cross-sectional evaluation of model outputs evaluates the accuracy, readability, transparency, objectivity, completeness, and bias of GPT-4o mini responses to common oculoplastic patient questions.MethodsThe study was reviewed and deemed exempt from institutional review board approval. Thirty-three questions were developed across eight clinical domains: floppy eyelid syndrome, cicatricial ectropion, entropion, blepharoplasty, facial aging, brow lift, ptosis, and facial nerve function. Seven domains contained four patient-centered questions, while the facial aging domain contained five questions, for a total of 33 questions. On September 16, 2024, GPT-4o mini was prompted for each question with a standardized prompt requesting responses from the perspective of an expert oculofacial surgeon, using the latest medical guidelines. Verbatim responses were recorded. Six American Society of Ophthalmic Plastic and Reconstructive Surgery fellowship-trained surgeons independently evaluated each response using a 5-point Likert scale assessing clinical accuracy, completeness, transparency, objectivity, and bias. Readability was assessed using established readability indices.ResultsGPT-4o mini demonstrated moderate clinical accuracy and completeness. Clinical accuracy was rated favorably in 51% of ratings (101/198), while completeness was rated favorably in 42% (84/198). Transparency and objectivity demonstrated stronger performance. 78% of ratings (155/198) agreed or strongly agreed that responses were transparent and 76% of ratings (150/198) agreed or strongly agreed that responses were objective. 8% of ratings (16/198) agreed or strongly agreed that Bias was present, while 65% (129/198) disagreed or strongly disagreed. Readability analyses demonstrated that ratings required college- to graduate-level reading ability.ConclusionGPT-4o mini generated transparent and objective ratings with low measured bias but demonstrated limitations in clinical accuracy, completeness, and readability. These findings apply specifically to GPT-4o mini and should not be generalized to other large language models. Future investigations should compare contemporary language models, evaluate alternative prompting strategies, and investigate multimodal systems that integrate text, clinical photographs, periocular anatomy, and patient symptoms.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.