Predicting outcomes in COVID-19: From internal validation to improving care
Bibliographic record
Abstract
One of the major challenges in treating patients with COVID-19 is predicting the severity of disease [[1]Richardson S. Hirsch J.S. Narasimhan M. Crawford J.M. McGinn T. Davidson K.W. et al.Presenting Characteristics, Comorbidities, and Outcomes Among 5700 Patients Hospitalized With COVID-19 in the New York City Area.JAMA. 2020; 323: 2052-2059Crossref PubMed Scopus (6118) Google Scholar]. In the context of a healthcare system stretched to capacity, the identification of factors associated with outcomes in COVID-19 is critically important [[2]Yadaw A.S. Li Y.C. Bose S. Iyengar R. Bunyavanich S. Pandey G Clinical features of COVID-19 mortality: development and validation of a clinical prediction model.Lancet Digit Health. 2020; 2: e516-ee25Summary Full Text Full Text PDF PubMed Scopus (161) Google Scholar]. Initially, COVID-19 was thought to be associated with a cytokine storm [[3]Feldmann M. Maini R.N. Woody J.N. Holgate S.T. Winter G. Rowland M. et al.Trials of anti-tumour necrosis factor therapy for COVID-19 are urgently needed.Lancet. 2020; 395: 1407-1409Summary Full Text Full Text PDF PubMed Scopus (422) Google Scholar]. Subsequently, cytokines and particularly IL-6 have attracted much attention as potential outcome biomarkers. However, it has been relatively difficult to consistently link poor clinical outcomes with baseline plasma concentrations of IL-6. A recent report indicates that IL-6 and IL-8 levels in critically ill patients with COVID-19 are considerably lower than in those with septic shock with or without the acute respiratory distress syndrome (ARDS) [[4]Kox M. Waalders N.J.B. Kooistra E.J. Gerretsen J. Pickkers P Cytokine Levels in Critically Ill Patients With COVID-19 and Other Conditions.JAMA. 2020; Crossref PubMed Scopus (228) Google Scholar]. Measures of individual cytokines at initial presentation may thus provide limited information on the clinical course of COVID-19. An alternative approach would be to study a composite cytokine profile and its change over the course of disease. In this issue of EBioMedicine, McElvaney and colleagues [[5]McElvaney O.J. Hobbs D.B. Qiao D. McElvaney O.F. Moll M. McEvoy N.L. et al.A linear prognostic score based on the ratio of interleukin-6 to interleukin-10 predicts outcomes in COVID-19.EBioMedicine. 2020; https://doi.org/10.1016/j.ebiom.2020.103026Summary Full Text Full Text PDF PubMed Scopus (54) Google Scholar] propose a new tool to aid in predicting the clinical outcome of patients hospitalized with COVID-19 based on the change observed over 4 days of the cytokine ratio of IL-6 to IL-10. The Dublin-Boston score is a 5-point scale, based on this change observed in 80 patients hospitalized with a confirmed diagnosis of COVID-19. Although baseline IL-6 was weakly associated with clinical outcome, the change in IL-6:IL-10 over time, particularly after 4 days, proved to be a far better predictor of outcome. In terms of prognostic research, this is a model development and internal validation study for the Dublin-Boston score [[6]Altman D.G. Vergouwe Y. Royston P. Moons K.G Prognosis and prognostic research: validating a prognostic model.BMJ. 2009; 338: b605Crossref PubMed Scopus (988) Google Scholar]. An exciting contribution of the study is identifying a new prognostic factor and demonstrating its superiority over IL-6. If the IL-6:IL-10 ratio is confirmed in a broader study population to have significant prognostic value, it can be widely used for prediction models related to COVID-19 and, perhaps, as a biomarker of treatment response. The predicted outcome of the Dublin-Boston score is a relative difference in clinical status between day 0 and 7, while the prediction is made using information between day 0 and 4. In order to demonstrate that it is potentially useful for clinicians, it is necessary that the predicted outcome not overlap known outcomes. Subsequent studies also need to demonstrate that the use of the model improves clinical decisions over not using the model [[7]Steyerberg E.W. Vickers A.J. Cook N.R. Gerds T. Gonen M. Obuchowski N. et al.Assessing the performance of prediction models: a framework for traditional and novel measures.Epidemiology. 2010; 21: 128-138Crossref PubMed Scopus (2948) Google Scholar]. The very determination of a “declined” or “improved” clinical status in the study hinged on the ability of the existing health care system to identify a change in status and make the appropriate decision to step up or step down care. Ideally, one should demonstrate that the model is [[1]Richardson S. Hirsch J.S. Narasimhan M. Crawford J.M. McGinn T. Davidson K.W. et al.Presenting Characteristics, Comorbidities, and Outcomes Among 5700 Patients Hospitalized With COVID-19 in the New York City Area.JAMA. 2020; 323: 2052-2059Crossref PubMed Scopus (6118) Google Scholar] better at making this assessment than usual parameters such as vital signs, mental status, kidney and liver function, and [[2]Yadaw A.S. Li Y.C. Bose S. Iyengar R. Bunyavanich S. Pandey G Clinical features of COVID-19 mortality: development and validation of a clinical prediction model.Lancet Digit Health. 2020; 2: e516-ee25Summary Full Text Full Text PDF PubMed Scopus (161) Google Scholar] that the clinical application of the model-based tool improves outcomes. It is important to emphasize that the ultimate indicator of a clinical prediction model's worth is its ability to impact care [[8]Moons K.G. Altman D.G. Vergouwe Y. Royston P Prognosis and prognostic research: application and impact of prognostic models in clinical practice.BMJ. 2009; 338: b606Crossref PubMed Scopus (653) Google Scholar]. Consistent with an internal validation study, the excellent performance characteristics of the model in this study should not be used to infer the potential usefulness of the model, only how it performs relative to other models in the same study set. The authors make the interesting comment that the IL-6:IL-10 ratio predicts outcomes but should not necessarily be used as a therapeutic target. There is an increasing awareness that prediction models should not only be predictive of the outcome, but should use predictors where the causal relationships with the outcome are more fully understood [[9]Prosperi M. Guo Y. Sperrin M. Koopman J.S. Min J.S. He X. et al.Causal inference and counterfactual prediciton in machine learning for actionable healthcare.Nature Machine Intel. 2020; 2: 369-375Crossref Scopus (111) Google Scholar]. If the IL-6:IL-10 ratio is not causally responsible for a change in status, then it is possible that unmeasured confounding factors, such as the administration of tocilizumab as highlighted by the authors, can change the observed value in a way that makes model predictions inaccurate. Two questions arise: First, what else has the potential to confound the predictive ability of the IL-6:IL-10 ratio in clinical practice? And second, if the IL-6:IL-10 ratio is not causally related, then what is? The first is the concern of a prediction model researcher who wants to make a useful model while the second is the concern of a basic science researcher who wants to explain the disease and find treatments, but both are fundamentally related. Finally, as more treatments move to the bedside, the Dublin-Boston score or the IL-6:IL-10 ratio could provide a useful tool to monitor response and support decisions to initiate or change therapies. Subsequent validation studies should include sensitivity analyses for patients on new therapies such as high-dose steroids, as is recently indicated for patients with severe and critical COVID-19 [[10]Lamontagne F. Agoritsas T. Macdonald H. Leo Y.S. Diaz J. Agarwal A. et al.A living WHO guideline on drugs for covid-19.BMJ. 2020; 370: m3379Crossref PubMed Scopus (465) Google Scholar]. Overall, the authors make a compelling case for studying composite cytokine profiles as biomarkers for patients with COVID-19. Hopefully, the initial promise of the IL-6:IL-10 ratio evolution will be pursued in subsequent studies to determine its impact on care. A linear prognostic score based on the ratio of interleukin-6 to interleukin-10 predicts outcomes in COVID-19The Dublin-Boston score is easily calculated and can be applied to a spectrum of hospitalized COVID-19 patients. More informed prognosis could help determine when to escalate care, institute or remove mechanical ventilation, or drive considerations for therapies. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".