Development and Validation of a Machine Learning Model to Estimate Risk of Adverse Outcomes Within 30 Days of Opioid Dispensation
Bibliographic record
Abstract
Importance: Machine learning approaches can assist opioid stewardship by identifying high-risk opioid prescribing for potential interventions. Objective: To develop a machine learning model for deployment that can estimate the risk of adverse outcomes within 30 days of an opioid dispensation as a potential component of prescription drug monitoring programs using access to real-world data. Design, Setting, and Participants: This prognostic study used population-level administrative health data to construct a machine learning model. This study took place in Alberta, Canada (from January 1, 2018, to December 31, 2019), and included all patients 18 years and older who received at least 1 opioid dispensation from a community pharmacy within the province. Exposures: Each opioid dispensation served as the unit of analysis. Main Outcomes and Measures: Opioid-related adverse outcomes were identified from administrative data sets. An XGBoost model was developed on 2018 data to estimate the risk of hospitalization, an emergency department visit, or mortality within 30 days of an opioid dispensation; validation on 2019 data was done to evaluate model performance. Model discrimination, calibration, and other relevant metrics are reported using daily and weekly predictions on both ranked predictions and predicted probability thresholds using all data from 2019. Results: A total of 853 324 participants represented 6 181 025 opioid dispensations, with 145 016 outcome events reported (2.3%); 46.4% of the participants were men and 53.6% were women, with a mean (SD) age of 49.1 (15.6) years for men and 51.0 (18.0) years for women. Of the outcome events, 77 326 (2.6% pretest probability) occurred within 30 days of a dispensation in the validation set (XGBoost C statistic, 0.82 [95% CI, 0.81-0.82]). The top 0.1 percentile of estimated risk had a positive likelihood ratio (LR) of 28.7, which translated to a posttest probability of 43.1%. In our simulations, the weekly measured predictions had higher positive LRs in both the highest-risk dispensations and percentiles of estimated risk compared with predictions measured daily. Net benefit analysis showed that using machine learning prediction may not add additional benefit over the entire range of probability thresholds. Conclusions and Relevance: These findings suggest that prescription drug monitoring programs can use machine learning classifiers to identify patients at risk of opioid-related adverse outcomes and intervene on high-risk ranked predictions. Better access to available administrative and clinical data could improve the prediction performance of machine learning classifiers and thus expand opioid stewardship efforts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".