Broadening the definition of patient-safety events: lessons from a multicentre learning health system collaborative
Bibliographic record
Abstract
BACKGROUND: Improving safety in healthcare has been paramount for decades, yet despite major attention and investment, improvement has remained incremental. Patient safety is a major concern in US healthcare, leading to significant harm and economic losses annually. Accurately identifying safety events remains difficult due to methodological discrepancies and lack of standardisation. This study evaluated the feasibility of implementing a standardised case-review methodology and safety-event taxonomy across diverse hospital settings to assess opportunities for improvement (OFIs) and compare findings with traditional definitions. METHODS: This multicentre retrospective cohort study reports data from 103 hospitals across the USA and Canada between 2016 and 2023. A multivariable logistic regression was performed to test case reviews for differences in the presence of one or more OFIs across several hospital types (bed size, academic status, urban setting, trauma level and Centres for Medicare and Medicaid Services overall star rating) and patient characteristics (age, gender, length of stay, admission and discharge code status and mortality). RESULTS: 19 181 cases were reviewed across the Learning Health System Collaborative, with a median of 107 reviews per hospital. Mortality was the most common cohort selection, studied by 91 hospitals (88%). At least one OFI was identified in 12 714 cases (66.3%). The logistic regression analysis found that all hospital characteristics and patient age, length of stay, code status and discharge disposition were significantly associated with at least one OFI. Of the 46 444 OFIs identified, 41 439 (89%) were from categories focused on omissions of care. The categories of end-of-life, documentation and treatment/care alone accounted for 25 980 OFIs (56%). CONCLUSION: The highest volumes of safety-related OFIs were associated with omissions of care, as opposed to the traditional definition of patient safety, which primarily includes outcomes from acts of commission.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".