What do end-users want to know about managing the performance of healthcare delivery systems? Co-designing a context-specific and practice-relevant research agenda
Bibliographic record
Abstract
BACKGROUND: Despite increasing interest in joint research priority-setting, few studies engage end-user groups in setting research priorities at the intersection of the healthcare and management disciplines. With health systems increasingly establishing performance management programmes to account for and incentivize performance, it is important to conduct research that is actionable by the end-users involved with or impacted by these programmes. The aim of this study was to co-design a research agenda on healthcare performance management with and for end-users in a specific jurisdictional and policy context. METHODS: We undertook a rapid review of the literature on healthcare performance management (n = 115) and conducted end-user interviews (n = 156) that included a quantitative ranking exercise to prioritize five directions for future research. The quantitative rankings were analysed using four methods: mean, median, frequency ranked first or second, and frequency ranked fifth. The interview transcripts were coded inductively and analysed thematically to identify common patterns across participant responses. RESULTS: Seventy-three individual and group interviews were conducted with 156 end-users representing diverse end-user groups, including administrators, clinicians and patients, among others. End-user groups prioritized different research directions based on their experiences and information needs. Despite this variation, the research direction on motivating performance improvement had the highest overall mean ranking and was most often ranked first or second and least often ranked fifth. The research direction was modified based on end-user feedback to include an explicit behaviour change lens and stronger consideration for the influence of context. CONCLUSIONS: Joint research priority-setting resulted in a practice-driven research agenda capable of generating results to inform policy and management practice in healthcare as well as contribute to the literature. The results suggest that end-users are keen to open the "black box" of performance management to explore more nuanced questions beyond "does performance management work?" End-users want to know how, when and why performance management contributes to behaviour change (or fails to) among front-line care providers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | no category Domain: not available · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | no category Domain: not available · Genre: Review About the Canadian research system: no · About a Canadian topic: no | Other design | low |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.144 | 0.015 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.009 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.006 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".