Development and validation of the ISARIC 4C Deterioration model for adults hospitalised with COVID-19: a prospective cohort study
Bibliographic record
Abstract
BACKGROUND: Prognostic models to predict the risk of clinical deterioration in acute COVID-19 cases are urgently required to inform clinical management decisions. METHODS: We developed and validated a multivariable logistic regression model for in-hospital clinical deterioration (defined as any requirement of ventilatory support or critical care, or death) among consecutively hospitalised adults with highly suspected or confirmed COVID-19 who were prospectively recruited to the International Severe Acute Respiratory and Emerging Infections Consortium Coronavirus Clinical Characterisation Consortium (ISARIC4C) study across 260 hospitals in England, Scotland, and Wales. Candidate predictors that were specified a priori were considered for inclusion in the model on the basis of previous prognostic scores and emerging literature describing routinely measured biomarkers associated with COVID-19 prognosis. We used internal-external cross-validation to evaluate discrimination, calibration, and clinical utility across eight National Health Service (NHS) regions in the development cohort. We further validated the final model in held-out data from an additional NHS region (London). FINDINGS: 74 944 participants (recruited between Feb 6 and Aug 26, 2020) were included, of whom 31 924 (43·2%) of 73 948 with available outcomes met the composite clinical deterioration outcome. In internal-external cross-validation in the development cohort of 66 705 participants, the selected model (comprising 11 predictors routinely measured at the point of hospital admission) showed consistent discrimination, calibration, and clinical utility across all eight NHS regions. In held-out data from London (n=8239), the model showed a similarly consistent performance (C-statistic 0·77 [95% CI 0·76 to 0·78]; calibration-in-the-large 0·00 [-0·05 to 0·05]); calibration slope 0·96 [0·91 to 1·01]), and greater net benefit than any other reproducible prognostic model. INTERPRETATION: The 4C Deterioration model has strong potential for clinical utility and generalisability to predict clinical deterioration and inform decision making among adults hospitalised with COVID-19. FUNDING: National Institute for Health Research (NIHR), UK Medical Research Council, Wellcome Trust, Department for International Development, Bill & Melinda Gates Foundation, EU Platform for European Preparedness Against (Re-)emerging Epidemics, NIHR Health Protection Research Unit (HPRU) in Emerging and Zoonotic Infections at University of Liverpool, NIHR HPRU in Respiratory Infections at Imperial College London.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".