Performance Management in Long-Term Care With Composite Measures of Quality
Bibliographic record
Abstract
Context: There is considerable interest in assessing quality of care in health care settings for performance management, improvement, accountability and recently also payment purposes. While individual providers and patients may want and need detailed indicators of quality performance specific to their needs, purchasers and regulatory bodies require aggregate or composite measures of quality that serve as robust signals of quality. Indicators must also be responsive to differences and changes in provider performance. Objectives: The objective of this paper is to evaluate (compare and contrast) hospital quality performance based on aggregations of quality indicators that are individually publicly reported for Long-Term Care (LTC) Hospitals in Ontario, Canada. Methods: We employed a retrospective facility-level cohort study. The cohort comprises 113 LTC hospitals from 1997 through 2005. We employed 12 valid and reliable risk-adjusted LTC quality indicators for this study representing both process and outcome quality measures. Adjusting for varying facility-level sample sizes, 95% binominal control limits were used to statistically identify significant differences in individual facility performance from provincial averages across each quality indicator. Aggregate measures were calculated using varying weights that adjusted for variability, reliability, correlation, and the health impact of different quality indicators. Observations: The prevalence of quality concerns ranged from 4% of residents (falls and worsening locomotion) to 31% (unregulated pain and worsening bladder continence). Most larger facilities tended to have more positive performance compared to moderate-sized and smaller facilities. Small sample sizes introduced considerable uncertainty in individual quality measures that was mitigated somewhat by aggregating individual quality measures. Aggregate indices were better able to mitigate small sample size concerns. While providing similar results, some variation in specific facilities identified as significantly higher and lower did exist across different weighting. Weighted composite indicators were able to identify LTC hospitals with particularly poor performance indicated by constituent quality indicators. Conclusions: An aggregate index of performance may be useful as an indicator of potentially underlying concerns in particular facilities. Transparent weighting mechanisms are useful to identify and emphasize priority areas for improvement. Priority weighting could be used to target pay-for-performance or other accountability initiatives. Individual facilities aiming to make improvements may still need specific information on individual indicators.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.037 | 0.104 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.007 | 0.011 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".