Home›Radiology & Imaging› AI-based screening tools demonstrate high sensitivity and specificity for obstructive sleep apnea detection
AI-based screening tools demonstrate high sensitivity and specificity for obstructive sleep apnea detectionAI tools show promise for screening obstructive sleep apnea
Journal of medical Internet researchPublished September 12, 2026Study authors: Lv Yujia, Zhou Lihui, Jiao Sihan, Lin Jiaying, Sun Boran, Hao Haixia, Jia Chenxiao, Wang Yuan, Bu Li…PubMed ↗DOI ↗Editorial oversight: Dr. Lars van Dijk, PhD · Surgical, Procedural & Diagnostic
AI-generated summary of the cited source, checked by automated accuracy review.
How we work
Share
Key Takeaway
Note that AI-based tools show high sensitivity and specificity for OSA screening, though evidence certainty is low.
This meta-analysis evaluated the diagnostic accuracy of AI-based screening tools for obstructive sleep apnea (OSA) using data from 60 studies. The analysis focused on sensitivity, specificity, and area under the receiver operating characteristic curve (AUC) at AHI thresholds of 5, 15, and 30 events per hour.
Key findings indicate high performance across all thresholds. Sensitivity was 0.94 at AHI 5, 0.87 at AHI 15, and 0.83 at AHI 30. Specificity was 0.77 at AHI 5, 0.81 at AHI 15, and 0.91 at AHI 30. Corresponding AUC values were 0.943, 0.907, and 0.920, respectively. The analysis also distinguished between non-PSG-derived tools and PSG-derived models, noting that PSG-derived models showed higher sensitivity (0.96, 0.90, 0.85) and specificity (0.82, 0.88, 0.96) at the respective AHI thresholds.
Authors noted substantial heterogeneity and limited external validation, resulting in low or very low certainty of evidence. Wide prediction intervals suggest variable performance across different populations. Clinically, these tools may assist in front-end screening and referral prioritization, while PSG-derived models may support reduced-channel assessment and sleep-laboratory workflows.
How this fits prior evidence
This meta-analysis addresses a gap in the clinical workflow for identifying patients with obstructive sleep apnea. While prior evidence establishes that OSA is associated with a 2- to 3-fold increase in stroke risk and is linked to worse glucose control in type 2 diabetes, this study focuses on the diagnostic accuracy of AI-based tools to identify such patients. The high sensitivity and specificity reported for AI-based screening tools may improve the identification of patients who could benefit from interventions like GLP-1 receptor agonists, which reduce AHI by 15.28 events per hour.
Living with obstructive sleep apnea can be exhausting, making it vital to catch the condition early. New research looked at how well artificial intelligence (AI) tools can screen for the condition. By looking at 60 different studies, researchers found that these AI tools performed well at identifying people with varying levels of sleep issues.
Specifically, the AI tools showed high sensitivity, which means they were very good at correctly identifying people who actually had the condition. For example, at a common clinical threshold, the tools showed a 0.94 sensitivity. While the evidence is still considered to have low certainty because the studies were very different from one another, the results suggest these tools could help doctors prioritize who needs a full sleep study.
These tools might be especially useful for initial screenings to help manage the flow of patients in a clinic. However, because the data comes from many different types of studies, the results might vary in different settings. Talk to your doctor about how these new screening methods might fit into your specific care plan.
What this means for you:
AI tools show high accuracy in identifying obstructive sleep apnea during initial screenings.
Common questions
How accurate are AI tools at finding sleep apnea?
AI tools showed high sensitivity, which means they are good at identifying people with the condition. For example, at a threshold of 5 events per hour, the sensitivity was 0.94. At other levels, sensitivity was 0.87 and 0.83. These results suggest AI can be a helpful tool for initial screening.
Can AI tools help doctors decide who needs a sleep study?
Yes, the data suggests that AI tools can help with front-end screening and prioritizing patients for referral. These tools can help manage the workflow in a sleep laboratory by identifying those who most need a full assessment.
Is the evidence for these AI tools very reliable yet?
The evidence is currently considered to have low or very low certainty. This is because the studies included were very different from one another and there was limited testing in different settings. You should discuss these options with your doctor.
BACKGROUND: Obstructive sleep apnea (OSA) is highly prevalent but remains substantially underdiagnosed. Polysomnography (PSG) is the reference standard, but its cost and limited availability constrain large-scale case identification. AI-based screening tools may support risk stratification and referral prioritization, but their diagnostic accuracy across apnea-hypopnea index (AHI) thresholds remains uncertain.
OBJECTIVE: This review aimed to systematically evaluate the diagnostic accuracy of AI-based OSA screening tools at AHI thresholds of ≥5, ≥15, and ≥30 events/hour, with emphasis on models using non-PSG-derived inputs.
METHODS: PubMed, Embase, Scopus, and Web of Science were searched for studies published from January 1, 2016, to May 3, 2026. Eligible studies included adults evaluated for suspected OSA or recruited from population-based cohorts, assessed AI-based models intended or interpretable for OSA screening, risk prediction, or screening-oriented severity classification, used PSG as the reference standard, and reported sufficient data to construct or reconstruct 2×2 contingency tables. Diagnostic accuracy was synthesized separately by AHI threshold and input source using bivariate random-effects models, with 95% CIs and prediction intervals (PIs). Risk of bias and certainty of evidence were assessed using QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies 2) and GRADE (Grading of Recommendations Assessment, Development, and Evaluation), respectively.
RESULTS: A total of 60 studies were included, of which 47 contributed data to the meta-analysis. At AHI thresholds of ≥5, ≥15, and ≥30 events/hour, pooled sensitivities were 0.94 (95% CI 0.92-0.96; 95% PI 0.71-0.99), 0.87 (95% CI 0.84-0.89; 95% PI 0.66-0.96), and 0.83 (95% CI 0.79-0.87; 95% PI 0.61-0.94), respectively; the corresponding specificities were 0.77 (95% CI 0.69-0.84; 95% PI 0.30-0.96), 0.81 (95% CI 0.75-0.85; 95% PI 0.39-0.96), and 0.91 (95% CI 0.87-0.94; 95% PI 0.55-0.99), respectively. The corresponding areas under the summary receiver operating characteristic curves were 0.943, 0.907, and 0.920. For non-PSG-derived tools, sensitivities were 0.92, 0.85, and 0.81, and specificities were 0.70, 0.74, and 0.85 at the 3 thresholds, respectively. For PSG-derived models, sensitivities were 0.96, 0.90, and 0.85, and specificities were 0.82, 0.88, and 0.96, respectively. Exploratory subgroup analyses suggested performance variation across selected study and model characteristics, including region, algorithmic framework, data source, and validation method.
CONCLUSIONS: AI-based tools showed generally favorable screening performance for OSA across clinically relevant AHI thresholds, although wide PIs suggest variable performance across future comparable populations and settings. By synthesizing diagnostic accuracy across 3 AHI thresholds and distinguishing non-PSG-derived from PSG-derived models, this review extends previous broad or modality-specific reviews and offers a clinically interpretable, pathway-specific basis for linking model performance to intended use. The findings may clarify potential roles for non-PSG-derived tools in front-end screening and referral prioritization and for PSG-derived models in reduced-channel assessment and sleep-laboratory workflow support. Given substantial heterogeneity, limited external validation, and low or very low certainty of evidence, prospective validation is needed before routine implementation.