Harnessing the Power of Technology to Transform Delirium Severity Measurement in the Intensive Care Unit: Protocol for a Prospective Cohort Study
Bibliographic record
Abstract
Background Delirium, an acute brain dysfunction, is a complication in up to 50% of patients in the intensive care unit (ICU). Measuring and mitigating delirium severity can reduce associated morbidity and improve long-term health outcomes post discharge. However, the perceived complexity of the available delirium detection tools and clinical workload limits the routine assessment of delirium severity. Developing a passive digital marker for delirium severity, combining routine electronic health record (EHR) and computer vision technology data, could be an implementable, scalable, and sustainable approach. Objective Our primary objective is to develop a passive digital marker for delirium severity (PDM-Del) and examine its performance in comparison to validated delirium severity tools. Our secondary objective is to evaluate the acceptability and usability of the PDM-Del by patients, families, and clinicians. Methods We will conduct a prospective, longitudinal cohort study to develop a PDM-Del using computer vision data and routinely collected EHR data. Following informed consent, the study team will collect image data through continuous digital video recordings of adult patients (>50 years) in their ICU room, routine EHR data (demographic and clinical variables), and administer delirium severity assessments (4 times daily) until ICU discharge or death. We will examine the usability and acceptability of the developed PDM-Del by patients, families, and direct care clinicians in a pilot randomized controlled clinical trial (aim 3). Descriptive statistics (means, SDs, medians, IQRs, and frequencies) and statistical differences between study instruments will be examined. We will use convolutional neural networks and machine learning to inform model development, testing, and validation. We will report model performance statistics, including accuracy, precision, recall, and the F1-score. Results We are currently in the recruitment and data collection phase. As of March 2025, we screened 3980 patients (32% eligible, n=1307), approached 665 (50%), and enrolled 150 participants (23% enrollment rate). Among the 150 patients, the median age was 67 (IQR 61-74) years, 62% (93/150) were male, and 91% (136/150) were White. Conclusions The PDM-Del could provide real-time, actionable feedback to direct care clinicians on the brain health of patients in the ICU. Early mitigation of delirium severity may decrease the risk of mortality, future Alzheimer disease and related dementia, and length of hospital stay. Trial Registration ClinicalTrials.gov NCT06172491; https://clinicaltrials.gov/study/NCT06172491 International Registered Report Identifier (IRRID) DERR1-10.2196/62912
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.066 | 0.055 |
| Meta-epidemiology (narrow) | 0.004 | 0.003 |
| Meta-epidemiology (broad) | 0.005 | 0.005 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.004 | 0.003 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.005 | 0.005 |
| Insufficient payload (model declined to judge) | 0.032 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".