Ranking hospital performance based on individual indicators: can we increase reliability by creating composite indicators?
Bibliographic record
Abstract
BACKGROUND: Report cards on the health care system increasingly report provider-specific performance on indicators that measure the quality of health care delivered. A natural reaction to the publishing of hospital-specific performance on a given indicator is to create 'league tables' that rank hospitals according to their performance. However, many indicators have been shown to have low to moderate rankability, meaning that they cannot be used to accurately rank hospitals. Our objective was to define conditions for improving the ability to rank hospitals by combining several binary indicators with low to moderate rankability. METHODS: Monte Carlo simulations to examine the rankability of composite ordinal indicators created by pooling three binary indicators with low to moderate rankability. We considered scenarios in which the prevalences of the three binary indicators were 0.05, 0.10, and 0.25 and the within-hospital correlation between these indicators varied between - 0.25 and 0.90. RESULTS: Creation of an ordinal indicator with high rankability was possible when the three component binary indicators were strongly correlated with one another (the within-hospital correlation in indicators was at least 0.5). When the binary indicators were independent or weakly correlated with one another (the within-hospital correlation in indicators was less than 0.5), the rankability of the composite ordinal indicator was often less than at least one of its binary components. The rankability of the composite indicator was most affected by the rankability of the most prevalent indicator and the magnitude of the within-hospital correlation between the indicators. CONCLUSIONS: Pooling highly-correlated binary indicators can result in a composite ordinal indicator with high rankability. Otherwise, the composite ordinal indicator may have lower rankability than some of its constituent components. It is recommended that binary indicators be combined to increase rankability only if they represent the same concept of quality of care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.058 | 0.094 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.007 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".