MétaCan
Menu
Back to cohort
Record W3085357481 · doi:10.1097/ccm.0000000000004630

Coronavirus Disease 2019 Prediction Modeling: Everything Old Is NEWS Again*

2020· letter· en· W3085357481 on OpenAlexaff
Stephanie Sibley, David M. Maslove

Bibliographic record

VenueCritical Care Medicine · 2020
Typeletter
Languageen
FieldMedicine
TopicSepsis Diagnosis and Treatment
Canadian institutionsQueen's University
Fundersnot available
KeywordsMedicinePandemicObservational studyEpidemiologyDiseaseCoronavirus disease 2019 (COVID-19)Health carePublic healthIntensive care medicineMEDLINEMedical emergencyInfectious disease (medical specialty)Pathology

Abstract

fetched live from OpenAlex

One of the most vexing challenges of the coronavirus disease 2019 (COVID-19) pandemic has been the incredible strain placed on healthcare resources—in particular, ICU resources—and its knock-on effects across healthcare systems worldwide. Pandemic planning and the development of contingencies to handle surges are largely the domain of public health and epidemiology, but they are informed by what we know about the likely clinical course of COVID-19 in various populations. This latter consideration is an exercise in clinical prediction modeling. It is understandable, therefore, that a great deal of attention has been paid to this issue over the last several months. A systematic review from Wynants et al (1) published in April of this year counted no fewer than 10 prognostic models for COVID-19. When the authors updated their search in June, the number of models had increased to 16. For the most part, these studies used clinical data from various epicenters of the pandemic in order to develop COVID-specific prediction models that estimate the probability of adverse outcomes, including progression to severe disease, and mortality. In this issue of Critical Care Medicine, Liu et al (2) add to this literature but take a slightly different approach. Rather than develop a de novo COVID-specific prediction model, the authors performed a retrospective, observational study to evaluate and compare the efficacy of several preexisting scoring systems for predicting in-hospital death in patients with COVID-19. Demographic data were collected from the charts of 673 consecutive patients admitted to the West Campus of the Wuhan Hospital with confirmed COVID-19 from January 30, 2020, to March 14, 2020, and the National Early Warning Score (NEWS), National Early Warning Score 2 (NEWS2), Rapid Emergency Medicine Score (REMS), CURB-65, and quick Sequential Organ Failure Assessment (qSOFA) were calculated for each patient. The authors found NEWS demonstrated the best discrimination for predicting in-hospital death with an area under the receiver operating characteristics curve (AUROC) of 0.882 (95% CI, 0.847–0.916). A NEWS score of greater than or equal to 5 was the optimal threshold, with a sensitivity of 84.3% and a specificity of 76.8%. NEWS2 and REMS also had good discrimination, and all the scores were well calibrated. The authors performed an analysis of the various subcomponents of the NEWS and found the oxygen saturation score alone also had good discrimination for prediction of in-hospital death, albeit with lower sensitivity than the complete NEWS. NEWS was developed to improve detection and response to clinical deterioration in adult patients with acute illness in hospital (3). It has also been deployed in the prehospital settings, albeit with weaker evidence of its utility (4). The score was updated in 2017 (NEWS2) with new indicators for the presence of hypercapnic respiratory failure and the use of supplemental oxygen (5). In contrast to REMS, CURB-65, and qSOFA scores, NEWS was not designed to predict mortality; its intended use was to monitor patients with repeated measurements over time, in order to detect clinical deterioration. The authors of the current study previously validated NEWS in a heterogeneous Chinese population of emergency intensive care patients (6), however, in that case, scores of greater than or equal to 7 were associated with an increased risk of death. Although the AUROC demonstrated in Liu et al (2) was similar to that of the general ICU study, the optimal cutoffs for making the prediction differed, suggesting the discrimination of this score may vary based on diagnosis and population. The National Institute for Health and Care Excellence note that the CURB-65 and NEWS2 have not been validated in patients with COVID-19 (7). The use of NEWS in COVID has previously been explored to a limited extent. Liao et al (8) used an adapted version of the NEWS that included age greater than or equal to 65 years as an independent risk factor with a point value of three to facilitate classification of illness severity and assist in admission decisions for patients with COVID-19. Patients were divided into four risk categories, with low-risk patients receiving routine monitoring and high-risk patients receiving continuous monitoring in a critical care setting. The study by Liao et al (8) focused on pandemic preparedness and no outcomes were reported. Hu et al (9) compared the Modified Early Warning Score (MEWS) to REMS for mortality prediction of critically ill patients with COVID-19 and found the REMS performed better than MEWS with an AUROC of 0.833. The models included in the review by Wynants et al (1) used patient characteristics that were associated with poor outcome from initial reports such as age, CT findings, sex, comorbidities, and serum markers such as d-dimers, C reactive protein, lactate dehydrogenase, and lymphocyte count. The intended use of these models was not well defined, the risk of bias was high, and the authors concluded that they could not be reliably used (1). Although Liu et al (2) have demonstrated the utility of using preexisting scores for prediction of mortality in patients with COVID-19, the study is not without limitations. It was conducted at a 800-bed hospital in Wuhan at the start of the coronavirus pandemic, where—by necessity—ICU level treatments such as ventilation and vasopressor infusions were at times provided outside of critical care areas. This may limit the generalizability of the findings to centers with a similar case mix and similar COVID-19 epidemiology. Furthermore, the oxygen saturation score essentially amounts to a new prediction model and should therefore be subjected to external validation. Undeniably, tools to predict mortality are important for guiding goals of care discussions and in the dreaded circumstance where triage rules are enacted in order to optimize resource allocation as equitably as possible. However, for clinicians admitting patients with COVID-19 to hospital, a more quotidian use case will be identifying those patients who are expected to have a stable clinical course and are suitable for a medical ward, as well as those more likely to deteriorate and require transfer to a higher level of care. NEWS and NEWS2 use repeated measurements of physiologic variables to anticipate decompensation and may be useful to optimize resource allocation and escalation of care when needed for patients with COVID-19. It is less clear if these scoring systems can inform the important question of which COVID-19 patients are likely to deteriorate and require a change in their level or care with single measurements at admission to the hospital. As COVID-19 continues to spread and resources are strained, there is a need for reliable, validated tools that allow clinicians to predict its clinical course. Although a bespoke COVID-19 prediction model may yield marginal improvements in discrimination, the AUROC value alone tells just one part of a larger story; even the most discriminant model is of little value if it is never used. With the rapid availability of clinical data for patients with COVID-19, an abundance of clinical decision rules is likely forthcoming. Few of these are likely to face robust external validation, and they may in fact add to the confusion already faced by medical providers (10). As the study by Liu et al (2) suggests, tools developed for general critical illness can be leveraged in the care of COVID-19 patients. The scores evaluated in the study are familiar to most and are already widely implemented in a variety of hospital settings. They may not be new, but they are tried and true, and as this study suggests, they are also COVID-ready.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.043
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.986
Threshold uncertainty score0.072

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.043
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0020.002
Science and technology studies0.0000.001
Scholarly communication0.0040.003
Open science0.0020.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.182
GPT teacher head0.398
Teacher spread0.216 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueCritical Care MedicineSame topicSepsis Diagnosis and TreatmentFrench-language works237,207