MétaCan
Menu
Back to cohort
Record W2953399894 · doi:10.1186/s12874-019-0769-x

Ranking hospital performance based on individual indicators: can we increase reliability by creating composite indicators?

2019· article· en· W2953399894 on OpenAlexfundno aff
Peter C. Austin, Iris Ceyisakar, Ewout W. Steyerberg, Hester F. Lingsma, Perla J. Marang‐van de Mheen

Bibliographic record

VenueBMC Medical Research Methodology · 2019
Typearticle
Languageen
FieldHealth Professions
TopicPatient Satisfaction in Healthcare
Canadian institutionsnot available
FundersCanadian Institutes of Health ResearchOntario Ministry of Health and Long-Term CareHeart and Stroke Foundation of Canada
KeywordsReliability (semiconductor)Ranking (information retrieval)Computer scienceReliability engineeringStatisticsPsychologyMedicineInformation retrievalMathematicsEngineering

Abstract

fetched live from OpenAlex

BACKGROUND: Report cards on the health care system increasingly report provider-specific performance on indicators that measure the quality of health care delivered. A natural reaction to the publishing of hospital-specific performance on a given indicator is to create 'league tables' that rank hospitals according to their performance. However, many indicators have been shown to have low to moderate rankability, meaning that they cannot be used to accurately rank hospitals. Our objective was to define conditions for improving the ability to rank hospitals by combining several binary indicators with low to moderate rankability. METHODS: Monte Carlo simulations to examine the rankability of composite ordinal indicators created by pooling three binary indicators with low to moderate rankability. We considered scenarios in which the prevalences of the three binary indicators were 0.05, 0.10, and 0.25 and the within-hospital correlation between these indicators varied between - 0.25 and 0.90. RESULTS: Creation of an ordinal indicator with high rankability was possible when the three component binary indicators were strongly correlated with one another (the within-hospital correlation in indicators was at least 0.5). When the binary indicators were independent or weakly correlated with one another (the within-hospital correlation in indicators was less than 0.5), the rankability of the composite ordinal indicator was often less than at least one of its binary components. The rankability of the composite indicator was most affected by the rankability of the most prevalent indicator and the magnitude of the within-hospital correlation between the indicators. CONCLUSIONS: Pooling highly-correlated binary indicators can result in a composite ordinal indicator with high rankability. Otherwise, the composite ordinal indicator may have lower rankability than some of its constituent components. It is recommended that binary indicators be combined to increase rankability only if they represent the same concept of quality of care.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.058
metaresearch head score (Gemma)0.094
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Science and technology studies, Research integrity, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch, Research integrity
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.060
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0580.094
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0020.002
Science and technology studies0.0020.001
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0010.007
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.280
GPT teacher head0.547
Teacher spread0.267 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations29
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueBMC Medical Research MethodologySame topicPatient Satisfaction in HealthcareFrench-language works237,207