Developing machine learning prescription anomaly detection models for opioid diversion surveillance in Ottawa capital region pharmacy chains
Bibliographic record
Abstract
Canada's opioid crisis has hit the Ottawa Capital Region hard, with prescription opioid diversion feeding street supply chains that contribute to overdose deaths. Current pharmacy-level diversion detection relies on pharmacist judgment and manual red-flag checklists, an approach that misses subtle patterns buried in high-volume dispensing data. This research developed and compared three machine learning models logistic regression, random forest, and XGBoost for automated detection of anomalous opioid prescriptions using dispensing records from 42 community pharmacies in Ottawa, Ontario. The dataset comprised 412,847 opioid dispensing events from January 2019 through December 2022. A panel of three clinical pharmacists labeled 1,847 confirmed or strongly suspected diversion cases as ground truth. Features included refill timing, prescriber-patient distance, dose escalation rate, prescriber diversity per patient, and geographic fill patterns. XGBoost achieved the highest area under the ROC curve (0.97), sensitivity (94.3%), and specificity (96.1%), outperforming random forest (AUC 0.94) and logistic regression (AUC 0.85). Early refill patterns were the most common anomaly type detected (34.7%), followed by doctor shopping (24.1%). Feature importance analysis identified days-to-early-refill and prescriber count per patient as the two strongest predictors. These findings support the deployment of XGBoost-based anomaly screening in Canadian community pharmacy dispensing software to assist pharmacists in identifying potential opioid diversion events in real time.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".