Development and Validation of the Sequential Organ Failure Assessment (SOFA)-2 Score
Bibliographic record
Abstract
Importance: Acute dysfunction of vital organs is the hallmark of critical illness. The Sequential Organ Failure Assessment (SOFA) score, the most widely adopted approach to describe organ dysfunction, has not been updated in 30 years and therefore may not appropriately capture current clinical practice and outcomes. Objectives: To inform the data-driven component of an updated score (SOFA-2) in varied geographical and resource settings (stages 6-8) after expert input via a modified Delphi process (stages 1-5). Design, Setting, and Participants: A federated analysis was performed on data collected from adult patients admitted to 1319 intensive care units (ICUs) in 9 countries (Australia, Austria, Brazil, France, Italy, Japan, Nepal, New Zealand, United States) between 2014 and 2023. Four representative multicenter cohorts containing data from 2 098 356 patients were used for data-driven score development and internal validation. External validation was performed on 6 cohorts containing data from 1 241 114 patients. Main Outcomes and Measures: Content validity for organ dysfunction identified through the modified Delphi process should be reflected by predictive validity using the area under the receiver operating characteristic (AUROC) curve of the score measured on the first ICU day (higher scores indicate worse organ dysfunction). Results: Of 3.34 million patient encounters, 270 108 (8.1%) died in the ICU (range, 4.5% to 20.5% across the 10 cohorts). SOFA-2 modified the 6 organ systems of the original SOFA score (brain, respiratory, cardiovascular, liver, kidney, hemostasis), including new variables and revised thresholds that better describe the organ dysfunction distribution from 0 to 4 points and their associated mortality (SOFA-2 AUROC, 0.79; 95% CI, 0.76-0.81; SOFA-1 AUROC, 0.77; 95% CI, 0.74-0.81). Evaluation of sequential SOFA-2 data from ICU day 1 to day 7 maintained its predictive validity. Insufficient data and lack of content validity precluded incorporation of gastrointestinal and immune dysfunction scores into SOFA-2. Conclusions and Relevance: The SOFA-2 score, updated to include contemporary organ support treatments and new score thresholds, describes organ dysfunction in a large, geographically and socioeconomically diverse population of critically ill adults.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Bench or experimental | high |
| gpt | no category Domain: not available · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | low |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".