Improving Administrative Code-Based Algorithms for Sepsis Surveillance*
Bibliographic record
Abstract
Sepsis is a global health priority and the focus of widescale efforts to improve diagnosis, treatment, and prevention (1). Accurate epidemiologic surveillance is critical to interpreting the impact of these efforts and informing future resource investments. Historically, most surveillance studies have used hospital discharge diagnosis codes to identify sepsis. This has clear advantages: administrative data are routinely generated for every hospitalization, readily available to researchers, relatively straightforward to analyze, and conducive to calculating nationwide or even global estimates that would otherwise be impossible to determine. However, administrative data have important limitations for sepsis surveillance. For one, multiple different combinations of codes have been used—broadly categorized as “implicit” definitions when they require both infection and organ dysfunction codes, or “explicit” definitions when they only use sepsis-specific codes—and prior research has demonstrated wide variability in sepsis incidence, outcomes, and accuracy on medical record review-based validations using different published definitions (2,3). More importantly, all administrative definitions are susceptible to changing diagnosis and coding practices over time as well as variability between hospitals and providers, complicating the interpretation of temporal trends and regional comparisons (4). These limitations have increasingly led researchers, public health officials, and policy makers to turn to electronic health record (EHR)-based clinical surveillance for sepsis. The most prominent example of this is the U.S. Centers for Disease Control and Prevention’s (CDC’s) Adult Sepsis Event (ASE) definition, which identifies patients with clinical indicators of presumed serious infection (blood cultures and ≥ 4 consecutive days of antibiotics) and concurrent acute organ dysfunction (vasopressors, mechanical ventilation, elevated lactate, or changes in baseline creatinine, bilirubin, or platelets) and has been applied to hundreds of hospitals with diverse EHRs to estimate the U.S. national burden of sepsis (5,6). Compared with administrative definitions, ASE generates more credible estimates of trends over time, displays less variability in incidence and mortality rates across hospitals, and has demonstrated comparable or better performance characteristics in medical record review-based validation studies (5,7–9). The major limitation of EHR-based sepsis surveillance is that EHR systems are not yet universally available, even in high-income countries. A survey of 27 high-income countries conducted in 2021, for example, found that only 15 had adopted or were in the process of adopting a unified EHR system (10). Even when EHRs are available, interoperability is often a challenge and implementing ASE surveillance requires informatics expertise and resources. As such, administrative code-based algorithms remain the only practical approach to sepsis surveillance in many countries for the foreseeable future, and improvements in this methodology are sorely needed. In this issue of Critical Care Medicine, Garland et al (11) report on an effort to develop an improved administrative definition of sepsis using a more transparent, systematic, and consensus-based approach compared with prior algorithms. Specifically, two expert panels consisting of experienced Canadian physicians with certification in critical care and/or infectious diseases evaluated 1928 infection and 108 organ dysfunction codes used in Canadian hospital abstracts and rated each using a 4-point Likert scale based on the likelihood that they could cause sepsis (for infection codes) or result from sepsis (for organ dysfunction codes). The experts underwent two phases of blinded voting to form consensus scores that were used to derive their primary algorithm (“AlgorithmL,” named for including codes with consensus scores of likely or higher) that ultimately included the two explicit sepsis codes, 48 organ dysfunction codes, and 720 infection codes. The investigators then applied their algorithm to population statistics and data from four hospitals in Calgary, Canada and compared their incidence and outcomes with four previously described International Classification of Diseases, 10th revision (ICD-10) code-based algorithms (one generated from a cross-walk of the International Classification of Diseases, 9th revision [ICD-9] based algorithm used by Angus et al [12] to ICD-10 codes; a narrow and broad version of the codes used by Fleischmann-Struzek et al [13]; and explicit severe sepsis or septic shock codes) as well as CDCs ASE definition. The primary findings are as follows. First, the number of codes used in the comparison algorithms varied from 42 to 941 infection codes and 2–36 organ dysfunction codes, and there was a high degree of nonoverlap with codes used in AlgorithmL. Second, AlgorithmL generated a sepsis incidence of 383 per 100,000 population in 2018 (6.4% of hospital admissions), with an in-hospital mortality rate of 18.7%. This represented the highest incidence and lowest mortality rate among the ICD-10 code-based algorithms and was comparable to the ASE definition. Third, sepsis incidence varied nearly ten-fold across algorithms, ranging from 40 of 100,000 to 383 of 100,000; hospital mortality rates also showed substantial variation, reaching as high as 40.8% with the explicit sepsis algorithm. The investigators conclude that AlgorithmL, by virtue of its inclusion of more infection and organ dysfunction codes compared with other code-based algorithms, captures milder cases of sepsis and may better encompass the full range of sepsis severity. There are several important contributions from the study by Garland et al (11). First and foremost, in contrast to prior studies where the rationale behind code selection has often been unclear, the investigators used a systematic and multidisciplinary expert consensus process to select codes, ensuring a robust and clinically relevant set of criteria that also provides a replicable framework for defining sepsis in administrative datasets in future clinical research. Second, their findings crystallize just how different various administrative definitions are with respect to the precise codes that are included in their algorithms, while reinforcing previous studies that have demonstrated wide variation in incidence and outcomes among these algorithms (2). Third, the finding that AlgorithmL (which purportedly captures “milder” forms of sepsis) still identified patients with an 18% in-hospital mortality rate reminds us that sepsis is a highly lethal disease that warrants even greater attention and resource investments. There are, of course, important limitations to the study by Garland et al (11), the most notable of which is the absence of any medical record review-based validation for AlgorithmL. The authors justify this by noting that that there is no true gold standard for sepsis diagnosis and even clinicians at the bedside struggle to identify sepsis because it is often unclear whether a patient is infected and whether organ dysfunction is due to a dysregulated host response to infection vs. a myriad of other potential processes (14). All of this is true, but previous studies have demonstrated that trained physicians can achieve adequate reliability in adjudicating sepsis by applying standardized definitions in medical record reviews (3,5). Furthermore, in our experience, while there are many “edge” cases where reasonable clinicians will disagree, medical record reviews can still shed important insight into the fidelity of sepsis algorithms by identifying unambiguous “false positive” cases where organ dysfunction is clearly unrelated to infection. This omission leaves room for uncertainty regarding AlgorithmL’s sensitivity and specificity, and, in particular, whether it is truly doing a better job at capturing the full range of sepsis severity or merely flagging more patients without bonafide sepsis. A second limitation is that the investigators did not consider the temporal relationship between infection and organ dysfunction codes, which could at least be approximated by the presence or absence of present-on-admission (POA) codes (i.e., POA infection codes and non-POA organ dysfunction codes, or vice versa, are unlikely to be related). In fairness, prior code-based sepsis algorithms, particularly in the ICD-9 era, also did not consider POA or non-POA status; however, POA coding has become more consistently used and accurate in recent years (15), and exploring the utility of POA flags in this setting could enhance the clinical relevance of administrative algorithms. Most importantly, even an optimized consensus set of codes is likely to suffer from the same problems that plague all administrative definitions of sepsis, namely temporal shifts and variability in diagnosis and coding practices. These limitations significantly impact the comparability of epidemiologic data derived from administrative codes across different time periods and geographical regions. In our view, parallel efforts are therefore needed to broaden the uptake of EHRs and EHR-based sepsis surveillance while continuing to use administrative definitions as an interim measure. Notably, although not the primary focus of the study by Garland et al (11), the inclusion of ASE in their analysis allows for an additional point of comparison to other studies: they found that ASEs hospital incidence and hospital mortality rate was 5.1% and 19.5%, respectively, which is very similar to numbers reported in the United States (5.9% and 15.6% [5]) and a large academic hospital in South Korea (5.4% and 16.6% [9]). This provides additional evidence for the reproducibility of ASE in identifying consistent sepsis cohorts in different high-income countries. In summary, the study by Garland et al (11) represents an important step forward in leveraging administrative data to measure the burden of sepsis. Although their new AlgorithmL will not overcome the numerous biases inherent to administrative data, a code-based approach remains the only realistic method for widescale sepsis surveillance across low-, middle-, and even many high-income countries for now. Additional refinement and validations of this new administrative definition across different countries may pave the way for more comprehensive research on global sepsis epidemiology while efforts to increase the adoption and feasibility of EHR-based clinical surveillance are ongoing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.203 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.009 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".