Pooled Cohort Probability Score for Subclinical Airflow Obstruction
Bibliographic record
Abstract
Abstract Rationale Early detection of chronic obstructive pulmonary disease (COPD) is a public health priority. Airflow obstruction is the single most important risk factor for adverse COPD outcomes, but spirometry is not routinely recommended for screening. Objectives To describe the burden of subclinical airflow obstruction (SAO) and to develop a probability score for SAO to inform potential detection and prevention programs. Methods Lung function and clinical data were harmonized and pooled across nine U.S. general population cohorts. Adults with respiratory symptoms, inhaler use, or prior diagnosis of COPD or asthma were excluded. A probability score for prevalent SAO (forced expiratory volume in 1 second/forced vital capacity < 0.70) was developed via hierarchical group-lasso regularization from clinical variables in strata of sex and smoking status, and its discriminative accuracy for SAO was assessed in the pooled cohort as well as in an external validation cohort (NHANES [National Health and Nutrition Examination Survey] 2011–2012). Incident hospitalizations and deaths due to COPD (respiratory events) were defined by adjudication or administrative criteria in four of nine cohorts. Results Of 33,546 participants (mean age 52 yr, 54% female, 44% non-Hispanic White), 4,424 (13.2%) had prevalent SAO. The incidence of respiratory events (N at-risk = 14,024) was threefold higher in participants with SAO versus those without (152 vs. 39 events/10,000 person-years). The probability score, which was based on six commonly available variables (age, sex, race and/or ethnicity, body mass index, smoking status, and smoking pack-years) was well calibrated and showed excellent discrimination in both the testing sample (C-statistic, 0.81; 95% confidence interval [CI], 0.80–0.82) and in NHANES (C-statistic, 0.83; 95% CI, 0.80–0.86). Among participants with predicted probabilities ⩾ 15%, 3.2 would need to undergo spirometry to detect one case of SAO. Conclusions Adults with SAO demonstrate excess respiratory hospitalization and mortality. A probability score for SAO using commonly available clinical risk factors may be suitable for targeting screening and primary prevention strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".