Screening Tools for Obstructive Sleep Apnea in Pregnant Women: An Extended and Updated Systematic Review and Meta-analysis
Bibliographic record
Abstract
The prevalence of obstructive sleep apnea syndrome (OSA) increases in women during pregnancy and negatively affects maternal and fetal outcomes. The updated systematic review and meta-analysis aimed to evaluate the validity of the Berlin, STOP-Bang, and Epworth sleepiness scale (ESS) questionnaires in detecting OSA in pregnant women. PubMed, Embase, and Web of Science were searched systematically up to March 2022. After eligible studies inclusion, two independent reviewers extracted demographic and clinical data. Bivariate random effects models were used to estimate the pooled accuracy measures including sensitivity and specificity, positive (PPV) and negative predictive values (NPVs), diagnostic odds ratio (DOR), and receiver operating characteristic curve (ROC) curve. We included 8 studies including 710 pregnant women with suspected OSA. The performance values of Berlin, STOP-Bang, and ESS questionnaires were as follows: the pooled sensitivity were 61% (95% confidence interval (CI): 40%-80%), 59% (95% CI: 49%-69%), and 29%, (95% CI: 10%-60%); pooled specificity were 61% (95% CI: 42%-78%), 80% (95% CI: 55%-93%), and 80% (95% CI: 50%-94%); pooled PPVs were 60% (95% CI: 0.49-0.72), 73% (95% CI: 61%-85%), and 59% (95% CI: 31%-87%); pooled NPVs were 60% (95% CI: 0.49-0.71), 65% (95% CI: 54%-76%), and 53% (95% CI: 41%-64%); and pooled DORs were 3 (95% CI: 1-5), 6 (95% CI: 2-19), and 2 (95% CI: 1-3), respectively. It seems that the Berlin, STOP-Bang, and ESS questionnaires had poor to moderate sensitivity and specificity in pregnancy, with the ESS showing the worst characteristics. Further studies are required to evaluate the performance of alternative screening methods for OSA in pregnancy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.009 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".