Mode
Text Size
Log in / Sign up

Structured workbook improves human and AI agreement during rehabilitation clinical practice guideline appraisalStructured Tools Improve Agreement Between Human Experts and AI Agents

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that structured workbooks improve agreement between human experts and AI agents during guideline appraisal.

This systematic review and meta-analysis evaluated 227 rehabilitation clinical practice guidelines (CPGs) to assess the impact of a structured appraisal workbook on agreement between human experts and AI agents. The study analyzed both English-language (163) and Chinese-language (64) guidelines. The analysis found that human expert agreement (ICC) improved from 0.66 to 0.84-0.92 across AGREE II domains when using the structured workbook. Similarly, AI models showed improved agreement when using the workbook, with DeepSeek-R1 increasing from 0.613 to 0.709 and o1-mini from 0.629 to 0.687.

Methodological quality was generally low, with applicability at 35.9% and stakeholder involvement at 52%. Notably, English-language guidelines outperformed Chinese-language guidelines in both scope and purpose (74.64 vs 68.88, P=.004) and applicability (39.14 vs 27.54, P<.001). Reporting quality via the RIGHT checklist was consistent (ICC=0.80-0.88), though reporting of funding and conflicts of interest was low at 44.1%.

AI agents completed appraisals significantly faster than humans (5.44 minutes vs 11.18 minutes). However, the authors emphasize that AI agents serve as efficient assistants rather than replacements for human oversight. Limitations include the specific focus on rehabilitation CPGs and the limited scope of English and Chinese languages. Further validation across other clinical specialties is required.

Researchers analyzed 227 rehabilitation clinical practice guidelines to see how well humans and AI agents could agree on their quality. The study compared experts working alone against those using a structured workbook. The results showed that using the workbook significantly improved the agreement between human experts and between different AI models.

While the AI agents were much faster at completing the evaluations than humans, the study notes that AI should be used as an assistant rather than a replacement for human experts. The study also found that while the tools improved agreement, they did not automatically guarantee that the guidelines themselves were of high quality.

This research is currently limited to rehabilitation guidelines in English and Chinese. Because the study only looked at one specific medical field, more research is needed across other medical specialties before these findings can be applied more broadly. Patients should know that these tools are intended to help experts work more efficiently, not to replace human judgment.

What this means for you:
Structured tools help humans and AI agree more on medical guidelines, but AI remains a tool for human experts.

Common questions

Can AI replace human experts in medical evaluations?

No, the study suggests that AI agents should serve as efficient assistants for human experts rather than replacing them. While AI agents were faster than humans, completing tasks in about 5.44 minutes compared to 11.18 minutes for humans, human oversight remains necessary for medical guideline appraisals.

Does using a structured workbook make the guidelines better?

Not necessarily. While the structured workbook improved the agreement between human experts and AI agents, it did not ensure that the guidelines themselves were of high quality. The study found that many guidelines still had low scores in areas like applicability and stakeholder involvement.

How did the AI perform compared to humans?

The study found that using a structured workbook improved agreement for both human experts and AI models. Specifically, the DeepSeek-R1 model saw its agreement score increase from 0.613 to 0.709, and the o1-mini model increased from 0.629 to 0.687 when using the structured guidance.

Study Details

Study typeMeta analysis
EvidenceLevel 1
PublishedAug 2026
View Original Abstract ↓
BACKGROUND: Rehabilitation clinical practice guidelines (CPGs) have increased rapidly, but inconsistent methodological quality limits their implementation. Although Appraisal of Guidelines for Research and Evaluation II (AGREE II) and Reporting Items for Practice Guidelines in Health Care (RIGHT) provide standardized appraisal frameworks, their application is time-consuming. Large language model (LLM)-based AI agents may offer a scalable alternative with uncertain reliability. OBJECTIVE: We evaluated rehabilitation CPGs' methodological and reporting quality and determined whether structured guidance improves human expert-AI agent agreement. METHODS: We systematically reviewed English- and Chinese-language rehabilitation CPGs from Embase, Scopus, PubMed, China National Knowledge Infrastructure, Wanfang Data, National Institute for Health and Care Excellence, Scottish Intercollegiate Guidelines Network, and Guidelines International Network up to June 2026. Methodological and reporting quality were assessed using AGREE II and the RIGHT checklist. Factors associated with guideline quality were examined using regression and subgroup analyses. Two AI agents were compared with human consensus with and without a structured guideline appraisal workbook, followed by external validation using 6 anterior cruciate ligament reconstruction CPGs. RESULTS: We included 227 CPGs (163 English-language, 64 Chinese-language). After introducing a structured guideline appraisal workbook, agreement among human experts improved markedly-mean intraclass correlation coefficients (ICCs) increased from -0.09 to 0.66 to 0.84-0.92 across AGREE II domains. Overall guideline quality remained low, with 35.9% (SD 18,8%) applicability and 52% (SD 17.2%) stakeholder involvement. English-language guidelines outperformed Chinese-language guidelines in scope and purpose (mean 74.64, SD 15.4 vs mean 68.88, SD 13.9; P=.004) and applicability (mean 39.14, SD 18.6 vs mean 27.54, SD 16.7; P<.001). Backward-elimination logistic regression revealed external review as an associated process characteristic (odds ratio 20.39, 95% CI 4.66-89.27; P<.001). RIGHT assessments showed consistent reliability (ICC=0.80-0.88). Reporting was highest for basic information (70.5%) and lowest for funding, declaration, and management of interests (44.1%). Meta-analysis of RIGHT reporting rates showed lower reporting among Chinese-language than English-language guidelines (risk difference [RD] -0.07, 95% CI -0.13 to -0.02, 95% prediction interval [PI] -0.39 to 0.24) and among guidelines published before vs after RIGHT release (RD -0.19, 95% CI -0.26 to -0.13, 95% PI -0.56 to 0.17). Without additional guidance, agent-human agreement was moderate (ICC=0.608-0.629). The workbook improved agreement for both models, with DeepSeek-R1's increasing from 0.613 to 0.709 and o1-mini's from 0.629 to 0.687. In validation beyond rehabilitation, DeepSeek-R1 maintained stable agreement (ICC=0.711) and completed appraisals in 5.44 minutes compared to 11.18 minutes for humans. CONCLUSIONS: Rehabilitation CPGs, particularly Chinese-language CPGs, continue showing deficiencies in applicability and stakeholder involvement. LLM-based appraisal without structured guidance provides insufficient agreement. Structured guidance improved agent-human agreement, supporting AI-assisted guideline appraisal under human oversight. Although further validation across additional clinical specialties is needed, AI agents can serve as efficient assistants in guideline appraisal instead of replacing humans. Future synthesis requires human-AI integration guided by structured, expert-defined principles. TRIAL REGISTRATION: PROSPERO CRD420251270676; https://www.crd.york.ac.uk/PROSPERO/view/CRD420251270676.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.