Validation of new ICD-10-based patient safety indicators for identification of in-hospital complications in surgical patients: a study of diagnostic accuracy
Bibliographic record
Abstract
OBJECTIVE: Administrative data systems are used to identify hospital-based patient safety events; few studies evaluate their accuracy. We assessed the accuracy of a new set of patient safety indicators (PSIs; designed to identify in hospital complications). STUDY DESIGN: Prospectively defined analysis of registry data (1 April 2010-29 February 2016) in a Canadian hospital network. Assignment of complications was by two methods independently. The National Surgical Quality Improvement Programme (NSQIP) database was the clinical reference standard (primary outcome=any in-hospital NSQIP complication); PSI clusters were assigned using International Classification of Disease (ICD-10) codes in the discharge abstract. Our primary analysis assessed the accuracy of any PSI condition compared with any complication in the NSQIP; secondary analysis evaluated accuracy of complication-specific PSIs. PATIENTS: All inpatient surgical cases captured in NSQIP data. ANALYSIS: We assessed the accuracy of PSIs (with NSQIP as reference standard) using positive and negative predictive values (PPV/NPV), as well as positive and negative likelihood ratios (±LR). RESULTS: We identified 12 898 linked episodes of care. Complications were identified by PSIs and NSQIP in 2415 (18.7%) and 2885 (22.4%) episodes, respectively. The presence of any PSI code had a PPV of 0.55 (95% CI 0.53 to 0.57) and NPV of 0.93 (95% CI 0.92 to 0.93); +LR 6.41 (95% CI 6.01 to 6.84) and -LR 0.40 (95% CI 0.37 to 0.42). Subgroup analyses (by surgery type and urgency) showed similar performance. Complication-specific PSIs had high NPVs (95% CI 0.92 to 0.99), but low to moderate PPVs (0.13-0.61). CONCLUSION: Validation of the ICD-10 PSI system suggests applicability as a first screening step, integrated with data from other sources, to produce an adverse event detection pathway that informs learning healthcare systems. However, accuracy was insufficient to directly identify or rule out individual-level complications.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.047 | 0.155 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".