MétaCan
Menu
← Back to cohort
Record W3034252606 · doi:10.1101/2020.06.14.150664

Machine Learning Identifies Complicated Sepsis Course and Subsequent Mortality Based on 20 Genes in Peripheral Blood Immune Cells at 24 Hours post ICU admission

2020· preprint· en· W3034252606 on OpenAlexaff
Shayantan Banerjee, Akram Mohammed, Hector R. Wong, Nades Palaniyar, Rishikesan Kamaleswaran

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2020
Typepreprint
Languageen
FieldMedicine
TopicSepsis Diagnosis and Treatment
Canadian institutionsHospital for Sick Children
FundersCincinnati Children's Hospital Medical Center
KeywordsSepsisDiseaseMedicineMachine learningPeripheral bloodIntensive care medicineInnate immune systemImmune systemArtificial intelligenceBioinformaticsImmunologyInternal medicineComputer scienceBiology

Abstract

fetched live from OpenAlex

Abstract A complicated clinical course for critically ill patients admitted to the ICU usually includes multiorgan dysfunction and subsequent death. Owning to the heterogeneity, complexity, and unpredictability of the disease progression, patient care is challenging. Identifying the predictors of complicated courses and subsequent mortality at the early stages of the disease and recognizing the trajectory of the disease from the vast array of longitudinal quantitative clinical data is difficult. Therefore, we attempted to identify novel early biomarkers and train the artificial intelligence systems to recognize the disease trajectories and subsequent clinical outcomes. Using the gene expression profile of peripheral blood cells obtained within 24 hours of PICU admission and numerous clinical data from 228 septic patients from pediatric ICU, we identified 20 differentially expressed genes that were predictive of complicated course outcomes and developed a new machine learning model. After 5-fold cross-validation with ten iterations, the overall mean area under the curve reached 0.82. Using the same set of genes, we further achieved an overall area under the curve of 0.72 when tested on an external validation set. This model was highly effective in identifying the clinical trajectories of the patients and mortality. Artificial intelligence systems identified eight out of twenty novel genetic markers SDC4, CLEC5A, TCN1, MS4A3, HCAR3, OLAH, PLCB1 and NLRP1 that help to predict sepsis severity or mortality. The discovery of eight novel genetic biomarkers related to the overactive innate immune system and neutrophils functions, and a new predictive machine learning method provides options to effectively recognize sepsis trajectories, modify real-time treatment options, improve prognosis, and patient survival. Research in Context Evidence before this study Transcriptomic biomarkers have long been explored as potential means of earlier disease endotyping. Much of the existing literature has however focused on mortality and discrete outcomes. Additionally, much of prior work in this area has been developed on statistical methods, while recent means of selecting features have not been sufficiently explored. Added value of this study In this study, we developed a robust machine learning based model for identifying novel biomarkers of complicated disease courses. We found 20 highly stable genes that predict disease complexity with an average derivation AUROC of 0.82 and validation AUROC of 0.72 within critically ill children, using peripheral blood collected within 24 hrs of ICU admission. Implications of all the available evidence Earlier identification of disease complexity can inform care management and targeted therapy. Therefore, the 20 gene candidates identified by our rigorous approach, can be used to identify, early in their ICU stay, patients who may ultimately develop significant organ dysfunction and complex care management.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.001
Threshold uncertainty score0.004

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.043
GPT teacher head0.281
Teacher spread0.238 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2020
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)→Same topicSepsis Diagnosis and Treatment→French-language works237,207→