Linking Electronic Health Record Prescribing Data and Pharmacy Dispensing Records to Identify Patient-Level Factors Associated With Psychotropic Medication Receipt: Retrospective Study
Bibliographic record
Abstract
Background: Pharmacoepidemiology studies using electronic health record (EHR) data typically rely on medication prescriptions to determine which patients have received a medication. However, such data do not affirmatively indicate whether these prescriptions have been filled. External dispensing databases can bridge this information gap; however, few established methods exist for linking EHR data and pharmacy dispensing records. Objective: We described a process for linking EHR prescribing data with pharmacy dispensing records from Surescripts. As a use case, we considered the prescriptions and resulting fills for psychotropic medications among pediatric patients. We evaluated how dispensing information affects identifying patients receiving prescribed medications and assessing the association between filling prescriptions and subsequent health behaviors. Methods: This retrospective study identified all new psychotropic prescriptions to patients younger than 18 years of age at Duke University Health System in 2021. We linked dispensing to prescribing data using proximate dates and matching codes between RxNorm concept unique identifiers and National Drug Codes. We described demographic, clinical, and service use characteristics to assess differences between patients who did versus did not fill prescriptions. We fit a least absolute shrinkage and selection operator (LASSO) regression model to evaluate the predictability of a fill. We then fit time-to-event models to assess the association between whether a patient filled a prescription and a future provider visit. Results: We identified 1254 pediatric patients with a new psychotropic prescription. In total, 976 (77.8%) patients filled their prescriptions within 30 days of their prescribing encounters. Thus, we set 30 days as a cut point for defining a valid prescription fill. Patients who filled prescriptions differed from those who did not in several key factors. Those who did not fill had slightly higher BMIs, lived in more disadvantaged neighborhoods, were more likely to have public insurance or self-pay, and included a higher proportion of male patients. Patients with prior well-child visits or prescriptions from primary care providers were more likely to fill. Additionally, patients with anxiety diagnoses and those prescribed selective serotonin reuptake inhibitors were more likely to fill prescriptions. The LASSO model achieved an area under the receiver operator characteristic curve of 0.816. The time to the follow-up visit with the same provider was censored at 90 days after the initial encounter. Patients who filled prescriptions showed higher levels of follow-up visits. The marginal hazard ratio of a follow-up visit with the same provider was 1.673 (95% CI 1.463-1.913) for patients who filled their prescriptions. Using the LASSO model as a propensity-based weight, we calculated the weighted hazard ratio as 1.447 (95% CI 1.257-1.665). Conclusions: Systematic differences existed between patients who did versus did not fill prescriptions. Incorporating external dispensing databases into EHR-based studies informs medication receipt and associated health outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.013 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.003 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".