Mode
Text Size
Log in / Sign up

AI burn depth assessment shows variable sensitivity and specificity across multiple imaging modalitiesTrial Shows AI May Help Assess Burn Depth Accuracy

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that AI burn depth assessment results are currently hypothesis-generating and not ready for clinical deployment.

This meta-analysis evaluates the diagnostic accuracy of AI-based burn depth assessment across various imaging modalities using a primary dataset of 2,541 observations and an expanded exploratory dataset of 4,897 observations. The analysis aimed to determine sensitivity and specificity for clinical use in burn management.

Primary dataset results were not pooled due to sparse 2x2 patterns and limited study numbers, with reported study-level sensitivity ranging from 0.50 to 1.00 and specificity from 0.54 to 0.97. The expanded exploratory dataset yielded a pooled sensitivity of 0.924 (95% CI 0.788-0.975) and a specificity of 0.877 (95% CI 0.701-0.956).

Several limitations impact the certainty of these findings, including mixed analytic units at the image, wound, and patient levels. Furthermore, evidence regarding skin tone impacts on specificity was sparse and non-confirmatory, and pediatric data were limited to a single stratum.

Because exploratory estimates rely on reconstructed 2x2 data, results are currently considered hypothesis-generating only. The authors conclude that the current evidence base is insufficient for the deployment-ready use of AI in burn depth assessment.

Researchers looked at 4,897 observations across several studies to see how well artificial intelligence (AI) identifies the depth of burns. They focused on whether AI could accurately measure how deep skin damage goes when looking at different types of medical images and wound evaluations.

The results showed that while some individual studies had high accuracy, the overall data was mixed. Because the researchers had to reconstruct parts of the data to get a clear picture, the findings are currently considered preliminary. The study also noted that there is very little information available regarding how these AI tools perform on different skin tones or in children.

Because the evidence is still limited and inconsistent, this technology is not yet ready for everyday use in hospitals. These results are meant to help researchers understand what is possible with AI. Patients should continue to rely on standard medical evaluations for burn care as current technology is still being tested.

What this means for you:
AI shows potential for assessing burn depth, but more research is needed before it can be used in clinical practice.

Common questions

Is this AI tool ready to be used in hospitals?

No, the evidence is not yet enough for it to be used in daily medical practice. The study found that the current data is inconsistent and limited. Experts say the results are only meant to help generate new ideas for future research rather than replacing current methods.

Does the AI work well on different skin tones?

The evidence regarding how the AI performs on different skin tones is currently sparse and not confirmed. Because of this lack of clear data, it is hard to say how accurately the tool works for everyone until more studies are completed.

How accurate was the AI in these studies?

In one specific group of data, the sensitivity was 0.924 and specificity was 0.877. However, other results varied widely between different studies. Because of these variations, the findings should be viewed as preliminary rather than a final proof of accuracy.

Study Details

Study typeMeta analysis
EvidenceLevel 1
PublishedJul 2026
View Original Abstract ↓
BACKGROUND: AI-based burn depth assessment is rapidly emerging, yet evidence for diagnostic accuracy, generalizability, and deployment readiness remains unclear. METHODS: We performed a PRISMA-DTA-aligned systematic review and meta-analysis. We synthesized accuracy outcomes and extracted deployment-relevant features, including validation design, reference-standard family, analytic unit, and subgroup reporting. Six studies (4,897 observations; mixed image-, wound-, and patient-level units) were included; four studies (N = 2541) formed the prespecified primary DTA dataset. We planned bivariate random-effects models where feasible and conducted exploratory subgroup analyses by imaging modality, skin tone, and age. Two additional studies were included only in exploratory analyses after prespecified approximate 2 × 2 reconstruction. RESULTS: In the primary dataset, the prespecified bivariate model did not converge because of sparse/extreme 2 × 2 patterns and limited study numbers; therefore, no pooled sensitivity or specificity was generated. Study-level sensitivity ranged from 0.50 to 1.00 and specificity from 0.54 to 0.97. In the expanded exploratory dataset (6 studies), pooled sensitivity was 0.924 (95% CI 0.788-0.975) and specificity 0.877 (95% CI 0.701-0.956), but these estimates are hypothesis-generating because they rely partly on approximately reconstructed 2 × 2 data. Exploratory descriptive analyses suggested possible modality-related variation and lower specificity in darker skin, although subgroup evidence was sparse and non-confirmatory; pediatric evidence was limited to a single within-study stratum. CONCLUSIONS: The evidence base is insufficient for deployment-ready use of AI burn depth assessment. The primary dataset did not support hierarchical pooling, and exploratory pooled estimates should be interpreted cautiously because they rely on reconstructed 2 × 2 data and mixed analytic units. More importantly, this review clarifies the evidence gaps separating promising algorithmic performance from deployable clinical decision support. Future studies should prioritize standardized reference standards, patient-level external validation in independent institutions, and reporting that supports safe integration into clinical workflows.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.