Mode
Text Size
Log in / Sign up

LLM-based risk-of-bias assessments show limited agreement with human reviewers in neurology prognosis studiesAI Tools Show Potential for Assessing Medical Research Quality

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that LLM-based risk-of-bias assessments currently show limited agreement with human reviewers in neurology studies.

This systematic review evaluates the utility of an LLM-based pipeline using zero-shot prompting to assess Risk-of-Bias (ROB) using the Quality in Prognosis Studies (QUIPS) framework. The review synthesized data from 298 articles across 15 reviews in the field of neurology, including conditions such as epilepsy, traumatic brain injury, and stroke.

Key findings indicate limited agreement between LLM and human reviewers, with a Cohen's weighted kappa of 0.22 (95% CI, 0.12 to 0.33). While Wilcoxon signed-rank tests showed statistically significant differences (p < 0.05) across four bias domains and overall risk scores, rank-biserial correlations indicated that human raters tended to assign higher risk scores than LLM counterparts.

A primary limitation noted is the small sample size (n=5) used to compare LLM-human versus human-human agreement, which suggests the LLM may not be inferior to human-human agreement but lacks robust statistical power. These findings suggest that while automated ROB assessments may reduce time and cost in systematic reviews, the current level of agreement between LLM and human experts remains limited.

How this fits prior evidence

This systematic review addresses a gap in methodology for systematic reviews in neurology. While previous coverage has focused on clinical interventions for stroke, such as tenecteplase and alteplase for acute ischemic stroke, or acupuncture for post-stroke foot drop, this study evaluates the technical feasibility of using LLMs to automate the assessment of study quality in the same clinical fields.

Researchers looked at how large language models (LLMs) perform when assessing the risk of bias in medical studies. They specifically looked at studies involving conditions like epilepsy, stroke, and traumatic brain injury. The study compared the scores given by AI to the scores given by human experts using a standard quality framework.

The results showed that while the AI and humans did not always agree perfectly, the AI might perform similarly to human-to-human comparisons in some cases. However, the study noted that the data for this specific comparison came from a very small sample size. The researchers also found that human experts tended to give higher risk scores than the AI did across several categories.

Because the study used a small sample for some comparisons, the results are not yet definitive. However, the findings suggest that automated tools could eventually save time and reduce costs for medical researchers. For now, these tools are seen as a potential way to speed up the review process rather than a replacement for human expertise.

What this means for you:
AI tools may help speed up research reviews, but they currently show different risk scores than human experts.

Common questions

How accurate is the AI compared to human experts?

The study found limited agreement between the AI and human reviewers. While the AI might not be inferior to human-to-human agreement in some cases, this finding was based on a very small sample of only 5 cases. Because of this small sample, the results are not yet conclusive.

Do the AI and humans give the same risk scores?

No, the results showed a significant difference in how scores were assigned. Human raters tended to assign higher risk scores than their AI counterparts across four different bias domains and in overall risk scores.

Can AI replace human experts in medical research?

The study suggests that automated tools could meaningfully reduce the time and cost of reviewing medical studies. However, because the AI and humans gave different scores, these tools are currently seen as a way to assist the process rather than a total replacement for human experts.

Study Details

Study typeSystematic review
Sample sizen = 5
EvidenceLevel 1
PublishedSep 2026
View Original Abstract ↓
Background: Risk-of-bias (ROB) assessments represent an integral component of systematic reviews. However, this task is often highly repetitive, time-consuming, and may lack inter-rater consistency. Large language models (LLMs) offer opportunities for automation in systematic reviews, which may expedite and enhance the quality and consistency of research synthesis. Methods: Using zero-shot prompting, we designed an LLM-based pipeline as a virtual mimic of a human reviewer for the Quality in Prognosis Studies (QUIPS) framework. Then, focusing on prognostic research in a single discipline (neurology), we applied this pipeline to articles included in previously published systematic reviews. We studied inter-rater agreement between both (1) the LLM and the original human ROB assessments and (2) between original human ROB assessments. Results: 298 articles from 15 reviews across three domains (epilepsy, traumatic brain injury, stroke) were included. We demonstrate the feasibility of a tailored, prompt-engineered LLM pipeline for automating ROB assessments with the QUIPS tool. While LLM-human agreement was limited (Cohen's weighted kappa; = 0.22, 95% CI, 0.12 - 0.33), our data tentatively suggest, based on a small sample (n=5), that it may not be inferior to human-human agreement (Cohen's weighted kappa; = -0.25, 95% CI, -1.04 - 0.54). Wilcoxon signed-rank tests were statistically significant (p < 0.05) across four bias domains and for the overall risk scores, and rank-biserial correlations demonstrated human raters' tendency to assign higher risk scores than LLM counterparts. Conclusions: With targeted methodological refinements - including standardization of QUIPS implementation and validation against expert ratings - automated ROB assessments may meaningfully reduce time and cost of systematic reviews of prognosis studies in neurology and beyond.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.