Mode
Text Size
Log in / Sign up

AI-generated videos for medical education currently lack standardized prompts and consistent quality assessment toolsAI Videos Face Quality Challenges in Medical Education Research

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note the lack of standardized prompts and low quality scores in currently available AI-generated medical education videos.

This scoping review examined the use of AI-generated videos for medical science popular media and education. The scope included an analysis of 10 studies to evaluate current practices, including prompt documentation, generation pathways, and evaluation instruments.

Key findings indicate a lack of standardization in the production process. Specifically, 5 out of 10 studies failed to clearly document the prompts used, with those that did provide only brief descriptions. Furthermore, AI-generated videos consistently received low Global Quality Scale (GQS) scores, with most scoring 2 out of 5. Only 1 out of 10 studies utilized three distinct evaluation tools, while others used fewer or none.

The authors note significant limitations including considerable heterogeneity among the included studies and a lack of standardized frameworks for prompts and assessment. The review suggests that current evidence is limited by the small sample size of 10 studies. These findings suggest that while AI-generated videos are being explored for education, the methodology for creating and evaluating these tools remains inconsistent.

A scoping review examined 10 studies to see how AI-generated videos are used to teach medical science. The researchers looked at the tools, prompts, and methods used to create these educational videos. They found that many of these projects lack a standard way to be measured or created.

One major finding was that half of the studies did not clearly document the specific instructions, or prompts, used to create the AI content. Additionally, when researchers graded the quality of these AI-generated videos using a standard scale, most received low scores of 2 out of 5. Only one study used three different tools to check for quality.

Because there is currently no standard way to make or judge these videos, it is hard to know if they are reliable for teaching. The evidence comes from a small number of studies and shows that more research is needed to create consistent rules for using AI in medical education.

What this means for you:
AI-generated medical videos currently show low quality scores and lack standardized methods for creation.

Common questions

Are AI-generated videos reliable for medical education?

Current evidence suggests caution. In the reviewed studies, most AI-generated videos received a low score of 2 out of 5 on a Global Quality Scale. Because many studies did not clearly document how they created the content, it is difficult to determine their reliability for teaching medical science at this time.

What are the main issues with using AI in medical education?

The main issues include a lack of standardized prompts and evaluation tools. Half of the studies reviewed failed to clearly document the instructions used to create the videos. This makes it hard for educators to ensure that the information provided is consistent or high-quality.

Study Details

Study typeSystematic review
EvidenceLevel 1
PublishedAug 2026
View Original Abstract ↓
ObjectivesThis scoping review aimed to systematically review research on the application of artificial intelligence-generated (AI-generated) videos for the general public education, patient education, medical student education, and physician education, and propose a framework covering the entire video production process.MethodsArksey and O’Malley’s scoping review guidelines were followed, using the PRISMA-ScR checklist. Studies were identified from PubMed, Web of Science, China National Knowledge Infrastructure and Google Scholar using keywords and medical subject headings terms related to “AI-generated”, “evaluation” and “video”, and screened for relevance, with any duplicates removed.ResultsThe initial 1,206 articles were reduced to 10 for review. Of these, half failed to clearly document the prompts used. Those that did provided only brief descriptions based on the application context or keywords, and the AI-generated videos consistently received low scores on validated instruments such as the Global Quality Scale (GQS), with most scoring 2 out of 5, indicating limited suitability for patient education. Comprehensive quality assessments were rare: only one study used three distinct evaluation tools, and the rest used one or two or none at all, revealing significant shortcomings in the standardization and comprehensiveness of the overall evaluation framework.ConclusionTo the best of our knowledge, this is among the first scoping reviews to synthesize the emerging evidence on AI-generated videos for the general public education, patient education, medical student education, and physician education. Few studies are reported and they demonstrate considerable heterogeneity. We identified insufficient standardization in prompts, generation pathways, tools, and evaluation instruments. We propose a procedural framework that may guide and inform future research in this field.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.