Validation and endorsement of health system performance measures for opioid use disorder in British Columbia, Canada: A Delphi panel study
Bibliographic record
Abstract
Background: Limited data exists on the performance of the healthcare system in opioid use disorder (OUD). We evaluated the face validity and potential risks of a set of health system performance measures for OUD collaboratively with clinicians, policymakers and people with lived experience of opioid use (PWLE) in the interest of establishing an endorsed set of measures for public reporting. Methods: Through a two-stage Delphi-panel approach, a panel of clinical and policy experts validated and considered 102 previously constructed OUD performance measures for endorsement using information on measurement construction, sensitivity analyses, quality of evidence, predictive validity, and feedback from local PWLE. We collected quantitative and qualitative survey responses from 49 clinicians and policymakers, and 11 PWLE. We conducted inductive and deductive thematic analysis to present qualitative responses. Results: A total of 37 measures of 102 were strongly endorsed (9/13 cascade of care, 2/27 clinical guideline compliance, 17/44 healthcare integration, and 9/18 healthcare utilization measures). Thematic analysis of responses revealed several themes regarding measurement validity, unintended consequences, and key contextual considerations. Overall, measures related to the cascade of care (excluding opioid agonist treatment dose tapering) received strong endorsements. PWLE highlighted barriers to accessing treatment, undignified aspects of treatment, and lack of a full continuum of care as their concerns. Conclusion: We defined 37 endorsed health system performance measures for OUD and presented a range of perspectives on their validity and use. These measures provide critical considerations for health system improvement in the care of people with OUD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.167 | 0.124 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.013 | 0.006 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.003 | 0.008 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".