MétaCan
Menu
Back to cohort
Record W4415666323 · doi:10.1001/jama.2025.20516

Development and Validation of the Sequential Organ Failure Assessment (SOFA)-2 Score

2025· article· en· W4415666323 on OpenAlexafffund
Otávio T. Ranzani, Mervyn Singer, Jorge I. Salluh, Manu Shankar‐Hari, David Pilcher, Joana Berger-Estilita, Craig M. Coopersmith, Nicole P. Juffermans, John G. Laffey, Matti Reinikainen, Ary Serpa Neto, Miguel Tavares, Jean-François Timsit, María del Pilar Arias López, Nish Arulkumaran, Diptesh Aryal, Elie Azoulay, Dipayan Chaudhuri, Dylan W. de Lange, Jan J. De Waele, Claúdia C. dos Santos, Bin Du, Sharon Einav, Teresa Engelbrecht, Fathima Fazla, Ricard Ferrer, Stefano Finazzi, Tomoko Fujii, Hayley B. Gershengorn, J Greene, Rashan Haniffa, Sicheng Hao, Mohd Shahnaz Hasan, Steve Hollenberg, Mariachiara Ippolito, Christian Jung, М. Yu. Кirov, Inès Lakbar, Jeffrey Lipman, Vincent X. Liu, Xiaoli Liu, Suzana M. Lobo, Demetrio Magatti, Greg S. Martin, Barbara Metnitz, Philipp Metnitz, Sheila Nainan Myatra, Simon Oczkowski, Fathima Paruk, Pirkka T. Pekkarinen, Lise Piquilloud, Anssi Pölkki, Hallie C. Prescott, Annika Reintam Blaser, Ederlon Rezende, Chiara Robba, Bram Rochwerg, Stéphane Ruckly, Rasoul Samei, Edward J. Schenck, Paul Secombe, Cornelius Sendagire, Moses Siaw-Frimpong, Andrew J. Simpkin, Márcio Soares, Charlotte Summers, Wojciech Szczeklik, Jukka Takala, Shiro Tanaka, Giovanni Tricella, Jean‐Louis Vincent, J. Wendon, Fernando G. Zampieri, Andrew Rhodes, Rui P. Moreno

Bibliographic record

VenueJAMA · 2025
Typearticle
Languageen
FieldMedicine
TopicSepsis Diagnosis and Treatment
Canadian institutionsUniversity of AlbertaImpactOccupational Cancer Research CentreUniversity of TorontoMcMaster UniversitySt. Joseph’s Healthcare Hamilton
FundersFaculdade de Ciências da Saúde, Universidade de MacauInstituto de Ciências Biomédicas Abel Salazar, Universidade do PortoMassachusetts Institute of TechnologyInsight SFI Research Centre for Data AnalyticsFundació Institut de Recerca Hospital Universitari Vall d’HebronMedizinische Universität GrazKarl-Franzens-Universität GrazUniversity College London Hospitals NHS Foundation TrustItä-Suomen YliopistoPeking Union Medical College HospitalUniversité de MontpellierUniwersytet Jagielloński Collegium MedicumMedizinische Universität WienJikei University School of MedicineSociedade Beneficente Israelita Brasileira Albert EinsteinUniversity of GalwayTartu ÜlikoolUniversity College LondonUniversité de LausanneChinese Academy of Medical SciencesUniversiteit GentHomi Bhabha National InstituteUniversity of BernKuopion Yliopistollinen SairaalaUniversity of AlbertaUniversity of TorontoHebrew University of JerusalemErasmus Medisch CentrumMonash UniversityInstitut National de la Santé et de la Recherche MédicaleKing's College LondonSchool of Medicine, Emory UniversityUniversitair Ziekenhuis GentPeking Union Medical CollegeEmory UniversityUniversität WienUniversidade da Beira InteriorWeill Cornell Medical CollegeHarvard T.H. Chan School of Public HealthIstituto di Ricerche Farmacologiche Mario Negri - IRCCSMcMaster UniversityUniversità degli Studi di PalermoKaiser PermanenteHelsingin YliopistoAssistance publique-Hôpitaux de ParisScience Foundation IrelandUniversity of PretoriaSt George's University Hospitals NHS Foundation TrustUniversity of MiamiUniversidade do Porto
KeywordsOrgan dysfunctionCritically illOrgan systemPopulationMEDLINERisk assessment

Abstract

fetched live from OpenAlex

Importance: Acute dysfunction of vital organs is the hallmark of critical illness. The Sequential Organ Failure Assessment (SOFA) score, the most widely adopted approach to describe organ dysfunction, has not been updated in 30 years and therefore may not appropriately capture current clinical practice and outcomes. Objectives: To inform the data-driven component of an updated score (SOFA-2) in varied geographical and resource settings (stages 6-8) after expert input via a modified Delphi process (stages 1-5). Design, Setting, and Participants: A federated analysis was performed on data collected from adult patients admitted to 1319 intensive care units (ICUs) in 9 countries (Australia, Austria, Brazil, France, Italy, Japan, Nepal, New Zealand, United States) between 2014 and 2023. Four representative multicenter cohorts containing data from 2 098 356 patients were used for data-driven score development and internal validation. External validation was performed on 6 cohorts containing data from 1 241 114 patients. Main Outcomes and Measures: Content validity for organ dysfunction identified through the modified Delphi process should be reflected by predictive validity using the area under the receiver operating characteristic (AUROC) curve of the score measured on the first ICU day (higher scores indicate worse organ dysfunction). Results: Of 3.34 million patient encounters, 270 108 (8.1%) died in the ICU (range, 4.5% to 20.5% across the 10 cohorts). SOFA-2 modified the 6 organ systems of the original SOFA score (brain, respiratory, cardiovascular, liver, kidney, hemostasis), including new variables and revised thresholds that better describe the organ dysfunction distribution from 0 to 4 points and their associated mortality (SOFA-2 AUROC, 0.79; 95% CI, 0.76-0.81; SOFA-1 AUROC, 0.77; 95% CI, 0.74-0.81). Evaluation of sequential SOFA-2 data from ICU day 1 to day 7 maintained its predictive validity. Insufficient data and lack of content validity precluded incorporation of gastrointestinal and immune dysfunction scores into SOFA-2. Conclusions and Relevance: The SOFA-2 score, updated to include contemporary organ support treatments and new score thresholds, describes organ dysfunction in a large, geographically and socioeconomically diverse population of critically ill adults.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmano category
Domain: not available · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Bench or experimentalhigh
gptno category
Domain: not available · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
models splitAgreement compares identical category sets and study designs across arms.

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.306
Threshold uncertainty score0.162

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.052
GPT teacher head0.340
Teacher spread0.287 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

The models applied no category: nothing in the taxonomy fit this work.

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designBench or experimental · Observational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations77
Published2025
Admission routes2
Has abstractyes

Explore more

Same venueJAMASame topicSepsis Diagnosis and TreatmentFrench-language works237,207