Multicenter Evaluation of an Interoperable System for Automated Guideline Adherence Monitoring in ICUs
Bibliographic record
Abstract
OBJECTIVE: To develop, apply, and validate a system for evaluating critical care guideline adherence, and to identify factors influencing real-world adherence across hospitals. DESIGN: Retrospective, multicenter observational study evaluating guideline adherence over 3.5 years and comparing automated adherence monitoring against expert human review. SETTING: Five university hospitals with different clinical information systems and data infrastructures. PATIENTS: A total of 82,000 intensive care episodes (2.2 million patient days). Six representative recommendations were selected from 41 intensive care guidelines and translated into a standardized digital format. Expert review encompassed more than 18,000 patient days. INTERVENTIONS: An automated system that applies digitally encoded guideline recommendations to standardized patient data extracted from hospital information systems. MEASUREMENTS AND MAIN RESULTS: The system determined, for each patient and recommendation, whether the recommendation applied (applicability) and whether treatment followed it (adherence). The primary outcome was the system's accuracy in identifying guideline applicability and adherence compared with manual clinician reviews. The secondary outcome was an analysis of how adherence to these recommendations varied and which factors influenced their real-world implementation. The system achieved 97.0% accuracy in identifying guideline applicability and adherence, significantly outperforming human reviewers (86.6% accuracy, p < 0.001; McNemar's test). The processing speed of the system exceeded 2000 patient days per second, compared with manual review at 2 patient days per minute. Adherence rates varied substantially across participating sites and over time, reflecting documentation inconsistencies, evolving clinical knowledge, and challenges in maintaining strict compliance. CONCLUSIONS: The guideline adherence monitoring system was successfully applied in multiple hospitals, demonstrating higher accuracy and efficiency compared with human review. Limitations of the system included dependence on consistent and structured documentation, as inconsistencies significantly complicate adherence monitoring. As the system is designed to support any guideline in the digital format used here, it provides a scalable solution for automated quality management in critical care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.047 | 0.077 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".