Predicting and Responding to Clinical Deterioration in Hospitalized Patients by Using Artificial Intelligence: Protocol for a Mixed Methods, Stepped Wedge Study
Bibliographic record
Abstract
BACKGROUND: The early identification of clinical deterioration in patients in hospital units can decrease mortality rates and improve other patient outcomes; yet, this remains a challenge in busy hospital settings. Artificial intelligence (AI), in the form of predictive models, is increasingly being explored for its potential to assist clinicians in predicting clinical deterioration. OBJECTIVE: Using the Systems Engineering Initiative for Patient Safety (SEIPS) 2.0 model, this study aims to assess whether an AI-enabled work system improves clinical outcomes, describe how the clinical deterioration index (CDI) predictive model and associated work processes are implemented, and define the emergent properties of the AI-enabled work system that mediate the observed clinical outcomes. METHODS: This study will use a mixed methods approach that is informed by the SEIPS 2.0 model to assess both processes and outcomes and focus on how physician-nurse clinical teams are affected by the presence of AI. The intervention will be implemented in hospital medicine units based on a modified stepped wedge design featuring three stages over 11 months-stage 0 represents a baseline period 10 months before the implementation of the intervention; stage 1 introduces the CDI predictions to physicians only and triggers a physician-driven workflow; and stage 2 introduces the CDI predictions to the multidisciplinary team, which includes physicians and nurses, and triggers a nurse-driven workflow. Quantitative data will be collected from the electronic health record for the clinical processes and outcomes. Interviews will be conducted with members of the multidisciplinary team to understand how the intervention changes the existing work system and processes. The SEIPS 2.0 model will provide an analytic framework for a mixed methods analysis. RESULTS: A pilot period for the study began in December 2020, and the results are expected in mid-2022. CONCLUSIONS: This protocol paper proposes an approach to evaluation that recognizes the importance of assessing both processes and outcomes to understand how a multifaceted AI-enabled intervention affects the complex team-based work of identifying and managing clinical deterioration. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/27532.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.046 | 0.050 |
| Meta-epidemiology (narrow) | 0.005 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.006 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.004 | 0.003 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.004 | 0.003 |
| Research integrity | 0.006 | 0.007 |
| Insufficient payload (model declined to judge) | 0.042 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".