Ranking hospitals according to acute myorcardial infarction mortality : do the methods matter?
Bibliographic record
Abstract
Background. Hospital performance indicators serve as a mechanism for making health care providers accountable to their patients. One indicator adopted by several jurisdictions is hospital mortality rates among patients with acute myocardial infarction (AMI). Despite potentially serious repercussions poor results can have on how a hospital is judged, there remains considerable variation in the methods used to measure and compare this indicator. The purpose of this study is to estimate the extent to which methods used to define AMI mortality outcomes and to deal with transferred AMI patients impact on hospital performance ratings. Methods. Using Quebec's Med-Echo hospital discharge records and vital statistics for 91,633 AMI patients admitted between 1992 and 1999, hospital rankings were compared using three methods to define AMI mortality outcome (in-hospital death, death within 7 days of admission, and death within 30 days of admission) and using three methods to handle transfers (excluding all transfers, including transfers while assigning the outcome to the initial hospital, and including transfers while assigning the outcome to the receiving hospital). Findings. There was discordance in hospital quintile classification 34% to 43% of the time when using pairwise comparisons of outcomes, and 23% to 32% of the time when using pairwise comparisons of ways to deal with transfers. Using hospital ranks to identify significant outliers as a method for evaluating hospitals, 5 hospitals were identified as "best performers" at least once, whereas 11 hospitals were identified "worst performers" at least once. One hospital was among the "worst performers" regardless of which among the six hierarchical analyses was used, while another was among the "best" using all but one analysis. The absolute difference in significantly high or low hospital mortality rates exceeded the clinically relevant benchmark of 1%. Conclusions. The methods used to define AMI mortality outcome, or to deal with transfers had an impact on which hospitals were identified as "outliers". Hospital reputations can be damaged by such findings. Furthermore, although this study was limited to comparing the impact on rankings based on AMI hospital mortality rates, other indicators of hospital performance may be influenced to a greater degree based on the methods used to deal with transferred patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.176 | 0.266 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.008 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.007 | 0.003 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".