Learning accurate personalized survival models for predicting hospital discharge and mortality of COVID-19 patients
Bibliographic record
Abstract
Since it emerged in December of 2019, COVID-19 has placed a huge burden on medical care in countries throughout the world, as it led to a huge number of hospitalizations and mortalities. Many medical centers were overloaded, as their intensive care units and auxiliary protection resources proved insufficient, which made the effective allocation of medical resources an urgent matter. This study describes learned survival prediction models that could help medical professionals make effective decisions regarding patient triage and resource allocation. We created multiple data subsets from a publicly available COVID-19 epidemiological dataset to evaluate the effectiveness of various combinations of covariates-age, sex, geographic location, and chronic disease status-in learning survival models (here, "Individual Survival Distributions"; ISDs) for hospital discharge and also for death events. We then supplemented our datasets with demographic and economic information to obtain potentially more accurate survival models. Our extensive experiments compared several ISD models, using various measures. These results show that the "gradient boosting Cox machine" algorithm outperformed the competing techniques, in terms of these performance evaluation metrics, for predicting both an individual's likelihood of hospital discharge and COVID-19 mortality. Our curated datasets and code base are available at our Github repository for reproducing the results reported in this paper and for supporting future research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".