Mode
Text Size
Log in / Sign up

LLM chatbots provide varying quality and clinical risk signals for vascular disease management informationChatbots show varying quality when giving advice on vascular disease

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that LLM chatbots provide inconsistent quality and clinical-risk signals for vascular disease information.

This guideline-based cross-sectional benchmark study evaluates the reliability of five publicly accessible LLM chatbot interfaces in providing information regarding vascular disease and perioperative management. The assessment utilized 22 guideline-derived English patient-facing questions to evaluate quality, readability, and potential clinical-risk signals. The study specifically analyzed responses using DISCERN, EQIP, GQS, and JAMA benchmark criteria.

Analysis of interrater agreement showed high consistency for DISCERN (ICC(A,1) = 0.940), EQIP (ICC(A,1) = 0.830), GQS (weighted kappa = 0.829), and JAMA benchmark criteria (weighted kappa = 0.898). However, the study found significant differences in performance across the various public-interface response sets for DISCERN, EQIP, and JAMA criteria (p < 0.05).

A noted limitation is that interface names were recorded as public-interface display labels rather than verified API-level model identifiers. These findings suggest that while LLMs can provide information on vascular disease, the quality and safety of the content vary significantly across different platforms. Clinicians should note that these results evaluate chatbot outputs rather than direct clinical outcomes in patients.

When patients search for information about vascular disease or surgery, they often turn to AI chatbots for quick answers. However, a new study shows that these tools do not provide consistent information. Researchers tested five different public AI interfaces by asking them 22 questions based on medical guidelines.

The study found significant differences in how the AI models performed across different scoring systems. While some models gave high-quality responses, others did not. This means a patient might get very different advice depending on which specific chatbot they happen to use.

It is important to remember that this study looked at the quality of the AI's text, not at actual patient outcomes. Because the AI responses vary so much, patients should be cautious when using these tools for medical information. Always talk to a doctor to get reliable information about vascular disease and surgery.

What this means for you:
AI chatbots provide inconsistent quality of information regarding vascular disease and surgical management.

Common questions

Can I rely on AI chatbots for information about vascular disease?

You should be cautious. This study found significant differences in the quality of information provided by different AI interfaces. Because the responses vary so much between different platforms, you should always consult a medical professional for reliable information regarding vascular disease and perioperative management.

How many different AI models were tested?

The study tested five different publicly accessible AI chatbot interfaces. These were evaluated using 22 questions based on medical guidelines to see how well they handled topics like vascular disease and surgery.

What did the study measure regarding the AI responses?

The study measured the quality, readability, and clinical-risk signals of the AI's answers. It used several scoring systems, including DISCERN, EQIP, and JAMA benchmark criteria, to see how well the AI followed medical guidelines.

Study Details

Study typeGuideline
EvidenceLevel 5
PublishedAug 2026
View Original Abstract ↓
BackgroundPatients increasingly use public large language model chatbot interfaces to seek health information. In vascular disease and perioperative management, default first responses may influence how patients interpret urgent symptoms, antithrombotic medications, procedural choices, and anesthesia-related safety issues.MethodsThis cross-sectional benchmark study evaluated 110 default first responses returned by five publicly accessible LLM chatbot interfaces during a defined access window on May 28–29, 2026, Beijing time (UTC + 8). Interface names were recorded solely as the public-interface display labels visible at the time of access and should not be interpreted as independently verified API-level model identifiers. Each model was queried with 22 guideline-derived English patient-facing questions, yielding 110 responses. Responses were assessed using DISCERN, Ensuring Quality Information for Patients (EQIP), Global Quality Scale (GQS), Journal of the American Medical Association (JAMA) benchmark criteria, used only as a visible metadata/transparency proxy, six readability formulas, an investigator-developed Guideline Concordance Score, and an investigator-developed Potential Clinical-Risk Severity Flag. Differences across the evaluated public-interface response sets were tested using Friedman tests with Holm-adjusted post hoc comparisons.ResultsInterrater agreement was high for established instruments: DISCERN ICC(A,1) = 0.940, EQIP ICC(A,1) = 0.830, GQS weighted κ = 0.829, and JAMA weighted κ = 0.898. Observed public-interface response performance differed significantly for DISCERN, EQIP, and JAMA criteria (all p 
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.