Mortality prediction in chronic obstructive pulmonary disease comparing the GOLD 2015 and GOLD 2019 staging systems: a pooled analysis of individual patient data
Bibliographic record
Abstract
In 2019, the Global Initiative for Chronic Obstructive Lung Disease (GOLD) proposed a clinical grading system for patients with chronic obstructive pulmonary disease (COPD) with 4 categories (A-D) based on symptoms and exacerbation history.As part of the COPD Cohorts Collaborative International Assessment (3CIA) initiative, we aim to compare mortality prediction of 2015 and 2019 COPD GOLD staging systems.We studied 17 139 COPD patients from the 3CIA study, selecting those with complete data. Patients were classified by the 2015 and 2019 GOLD ABCD systems and we compared the predictive ability for 5-year mortality of both classifications.17 139 patients with COPD were enrolled in 22 cohorts of 11 countries; 8823 of them had complete data and were analyzed. Mean age was 63.9 years(SD 9.8); 5 552(62.9%) were male and mean FEV1 was 54.8%(SD 22.3). Compared with 2015, the GOLD 2019 classified the patients in milder degrees of COPD:groups C and D decreased from 13.6% to 5.8% and 40% to 17.8% respectively (p< 0.001). Five-year mortality did not differ between groups B and C in GOLD 2015; in contrast, in GOLD 2019, five-year mortality was greater for group B than C.When analyzing individually the prognostic accuracy of each of the groups, patients classified as group A and B had better sensitivity and positive predictive value with GOLD 2019 classification than GOLD 2015. On the contrary, GOLD 2015 had better sensitivity for group C and D than GOLD 2019. The AUC for 5-year mortality were only 0.67 (95% CI 0.66-0.68) for GOLD 2015 and 0.65 (95% CI 0.63-0.66) for GOLD 2019.GOLD 2019 does not predict mortality better than GOLD 2015 system.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".