Application of patient safety indicators internationally: a pilot study among seven countries
Bibliographic record
Abstract
OBJECTIVE: To explore the potential for international comparison of patient safety as part of the Health Care Quality Indicators project of the Organization for Economic Co-operation and Development (OECD) by evaluating patient safety indicators originally published by the US Agency for Healthcare Research and Quality (AHRQ). DESIGN: A retrospective cross-sectional study. SETTING: Acute care hospitals in the USA, UK, Sweden, Spain, Germany, Canada and Australia in 2004 and 2005/2006. DATA SOURCES: Routine hospitalization-related administrative data from seven countries were analyzed. Using algorithms adapted to the diagnosis and procedure coding systems in place in each country, authorities in each of the participating countries reported summaries of the distribution of hospital-level and overall (national) rates for each AHRQ Patient Safety Indicator to the OECD project secretariat. RESULTS: Each country's vector of national indicator rates and the vector of American patient safety indicators rates published by AHRQ (and re-estimated as part of this study) were highly correlated (0.821-0.966). However, there was substantial systematic variation in rates across countries. CONCLUSIONS: This pilot study reveals that AHRQ Patient Safety Indicators can be applied to international hospital data. However, the analyses suggest that certain indicators (e.g. 'birth trauma', 'complications of anesthesia') may be too unreliable for international comparisons. Data quality varies across countries; undercoding may be a systematic problem in some countries. Efforts at international harmonization of hospital discharge data sets as well as improved accuracy of documentation should facilitate future comparative analyses of routine databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.019 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".