Development and validation of PRECISE-X model: predicting first severe exacerbation in COPD
Bibliographic record
Abstract
OBJECTIVES: In patients with chronic obstructive pulmonary disease (COPD), severe exacerbations (ECOPDs) impose significant morbidity and mortality. Current guidelines emphasise using ECOPD history to inform preventive treatments but offer limited guidance for risk stratification for the first severe ECOPD. METHODS: We developed and validated PRECISE-X using a cohort of newly diagnosed COPD patients from the UK's Clinical Practice Research Datalink (2004-2022), to predict first severe ECOPD over 5 years (primary outcome) and 12 months (secondary outcome). Predictors were selected via clinical expertise and data-driven methods. Internal-external cross-validation was performed across practice regions to evaluate the model's out-of-sample performance in terms of discrimination (c-statistic), calibration and net benefit. RESULTS: The study included 2 19 015 patients (mean age 66.0; 42.4% female). Observed risk of first severe ECOPD was 29.5% at 5 years (4.2% at 1 year). The final model included four mandatory predictors (sex, age, Medical Research Council dyspnoea score and forced expiratory volume in 1 second) and 28 optional predictors. In internal-external cross-validation, the average out-of-sample c-statistic was 0.836 (95% CI 0.827 to 0.846) for 5-year prediction and 0.756 (95% CI 0.746 to 0.766) for 1-year prediction. Calibration across regions was robust, and the model showed positive NB across a wide range of risk thresholds. In a secondary validation assessment among those with available spirometry data with confirmed airflow obstruction, the model was well calibrated and had only a modest decline in discriminatory performance. CONCLUSIONS: PRECISE-X accurately predicts the first severe COPD exacerbation using routine clinical data, supporting earlier risk stratification and proactive disease management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".