Researchers looked at how large language models (LLMs) perform when assessing the risk of bias in medical studies. They specifically looked at studies involving conditions like epilepsy, stroke, and traumatic brain injury. The study compared the scores given by AI to the scores given by human experts using a standard quality framework.
The results showed that while the AI and humans did not always agree perfectly, the AI might perform similarly to human-to-human comparisons in some cases. However, the study noted that the data for this specific comparison came from a very small sample size. The researchers also found that human experts tended to give higher risk scores than the AI did across several categories.
Because the study used a small sample for some comparisons, the results are not yet definitive. However, the findings suggest that automated tools could eventually save time and reduce costs for medical researchers. For now, these tools are seen as a potential way to speed up the review process rather than a replacement for human expertise.