Mode
Text Size
Log in / Sign up

Prediction models for prolonged air leak show variable discriminatory power ranging from 0.644 to 0.914Prediction models for lung surgery complications show mixed accuracy

AI-generated summary of the cited source, checked by automated accuracy review. How we work

Key Takeaway
Note that existing prediction models for prolonged air leak show significant heterogeneity and high risk of bias.

This systematic review evaluated prediction models designed to identify patients at risk for prolonged air leak (PAL) following pulmonary resection. The scope included an assessment of discriminatory power, calibration, and risk of bias across 26 different models.

The primary finding was a wide range in discriminatory power (AUC), with values reported between 0.644 and 0.914. These results reflect the inclusion of various models with differing components such as pleural adhesions, FEV1, BMI, age, smoking history, gender, and other factors.

The authors noted significant limitations including high overall risk of bias, single-center development, insufficient external validation, and incomplete calibration reporting. Substantial heterogeneity among the 26 models prevents a unified summary estimate of performance.

Clinical utility is currently limited by these methodological issues and variability in model performance. The review suggests that multicenter prospective studies are necessary to improve the reliability and clinical application of prediction tools for pulmonary surgery patients.

When a patient undergoes a pulmonary resection, a prolonged air leak can cause significant complications. Doctors use prediction models to try and identify which patients are at the highest risk for these issues. However, a review of 26 different models shows that their accuracy varies quite a bit.

The study found that the ability of these models to correctly identify a long air leak ranges from 0.644 to 0.914 on a scale used to measure predictive power. Because each model was developed differently and many were created at only one center, it is hard to say which ones are the most reliable for everyday use.

There are several hurdles to overcome before these tools can be used reliably in every hospital. Many models lacked outside testing or had issues with how they were reported. Because of this high level of variation and some flaws in how the data was collected, experts suggest that more large-scale studies across multiple hospitals are needed.

What this means for you:
Current prediction tools for lung surgery complications vary greatly in accuracy and need more testing to be reliable.

Common questions

How accurate are the current models for predicting air leaks?

The study looked at 26 different prediction models. These models showed a wide range of accuracy, with scores between 0.644 and 0.914. Because there is so much variation between the different models, it is difficult to say how well any single one will perform in a clinical setting.

What factors are used to predict these complications?

The prediction models look at several specific factors to identify risk. These include things like the patient's age, gender, body mass index (BMI), smoking history, and lung function measurements like FEV1.

Why aren't these models used more consistently yet?

Many of the current models were developed at only one center and have not been tested in different locations. There are also issues with how they were reported and a high risk of bias, meaning more large-scale studies across multiple hospitals are needed to make them reliable.

Study Details

Study typeMeta analysis
EvidenceLevel 1
PublishedAug 2026
View Original Abstract ↓
This systematic review aims to comprehensively map, critically appraise, and synthesize the quality and performance of existing prediction models for prolonged air leak (PAL) following pulmonary resection, with a focus on their clinical utility and potential for translation. A systematic search was conducted in CNKI, Wanfang Data, VIP, SinoMed, PubMed, Web of Science, Embase, and the Cochrane Library, up to March 2026. Two researchers independently screened the literature, extracted data, and assessed the risk of bias in the predictive models. A qualitative, narrative synthesis of model characteristics and performance was performed. Results: A total of 26 models were included. Most studies had good applicability, but the overall risk of bias was high. The models showed a wide range of discriminatory power, with Area Under the Curve(AUC)values ranging from 0.644 to 0.914. Frequently identified predictors included pleural adhesions, forced expiratory volume in 1 second(FEV1), body mass index(BMI), age, smoking history, and gender. Existing predictive models exhibit considerable variability in reported discriminatory performance, with AUCs ranging from 0.644 to 0.914. However, the majority of these models are hampered by issues such as single-center development, insufficient external validation, incomplete calibration reporting, and methodological biases. The substantial heterogeneity observed across the reviewed studies precludes the generation of a single, reliable summary estimate of model performance. To enhance the clinical applicability and practical value of these predictive tools, future research should prioritize multicenter prospective studies, optimize variable handling and model validation strategies, and ensure comprehensive reporting of both discrimination and calibration. https://www.crd.york.ac.uk/PROSPERO/home, identifier CRD420261435579.
Free Newsletter

Clinical research that matters. Delivered to your inbox.

Join thousands of clinicians and researchers. No spam, unsubscribe anytime.