MétaCan
Menu
Back to cohort
Record W4411016362 · doi:10.1097/cce.0000000000001275

Mortality Prediction Performance Under Geographical, Temporal, and COVID-19 Pandemic Dataset Shift: External Validation of the Global Open-Source Severity of Illness Score Model

2025· article· en· W4411016362 on OpenAlexaff
Takeshi Tohyama, Liam G. McCoy, Euma Ishii, Sahil Sood, Jesse D. Raffa, Takahiro Kinoshita, Leo Anthony Celi, Satoru Hashimoto

Bibliographic record

VenueCritical Care Explorations · 2025
Typearticle
Languageen
FieldMedicine
TopicSepsis Diagnosis and Treatment
Canadian institutionsUniversity of Alberta
Fundersnot available
KeywordsGeneralizability theoryMedicineStandardized mortality ratioCohortPandemicHealth careAPACHE IIEmergency medicineCohort studyIntensive careReceiver operating characteristicPredictive modellingCoronavirus disease 2019 (COVID-19)Intensive care unitIntensive care medicineDiseaseInternal medicineMachine learningStatisticsComputer scienceInfectious disease (medical specialty)

Abstract

fetched live from OpenAlex

BACKGROUND: Risk-prediction models are widely used for quality of care evaluations, resource management, and patient stratification in research. While established models have long been used for risk prediction, healthcare has evolved significantly, and the optimal model must be selected for evaluation in line with contemporary healthcare settings and regional considerations. OBJECTIVES: To evaluate the geographic and temporal generalizability of the models for mortality prediction in ICUs through external validation in Japan. DERIVATION COHORT: Not applicable. VALIDATION COHORT: The care Japanese Intensive care PAtient Database from 2015 to 2022. PREDICTION MODEL: The Global Open-Source Severity of Illness Score (GOSSIS-1), a modern risk model utilizing machine learning approaches, was compared with conventional models-the Acute Physiology and Chronic Health Evaluation (APACHE-II and APACHE-III)-and a locally calibrated model, the Japan Risk of Death (JROD). RESULTS: Despite the demographic and clinical differences of the validation cohort, GOSSIS-1 maintained strong discrimination, achieving an area under the curve of 0.908, comparable to APACHE-III (0.908) and JROD (0.910). It also exhibited superior calibration, achieving a standardized mortality ratio (SMR) of 0.89 (95% CI, 0.88-0.90), significantly outperforming APACHE-II (SMR, 0.39; 95% CI, 0.39-0.40) and APACHE-III (SMR, 0.46; 95% CI, 0.46-0.47), and demonstrating a performance close to that of JROD (SMR, 0.97; 95% CI, 0.96-0.99). However, performance varied significantly across disease categories, with suboptimal calibration for neurologic conditions and trauma. While the model showed temporal stability from 2015 to 2019, performance deteriorated during the COVID-19 pandemic, broadly reducing performance across disease categories in 2020. This trend was particularly pronounced in GOSSIS compared with APACHE-III. CONCLUSIONS: GOSSIS-1 demonstrates robust discrimination despite substantial geographic dataset shift but shows important calibration variations across disease categories. In particular, in a complex model like GOSSIS-1, stresses on the health system, such as a pandemic, can manifest changes in model calibration.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.016
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.016
Threshold uncertainty score0.074

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0140.016
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0010.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.206
GPT teacher head0.431
Teacher spread0.225 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueCritical Care ExplorationsSame topicSepsis Diagnosis and TreatmentFrench-language works237,207