Cross-validation of Drug Use Records in Two Pharmaceutical Databases: A Population-based Study of Alberta’s Tomorrow Project Cohort
Bibliographic record
Abstract
IntroductionPharmaceutical Information Network (PIN, 2008-now) is a provincial database collecting patients’ medication information in Alberta, Canada. Alberta Blue Cross (ABC), the largest health benefit provider in Alberta, has been managing pharmaceutical records for senior patients (65+ years) whose medications are covered by Alberta’s government-sponsored health benefit plan since 1970s.
 Objectives and ApproachOver 96% of participants in Alberta’s Tomorrow Project (ATP), a province-wide cohort study of cancer and chronic diseases in Canada, consented to data linkage to healthcare databases. To cross-validate medication records in the two pharmaceutical databases in Alberta, individual-level data of ATP participants aged 65+ years were cross-linked between PIN and ABC databases (2008-2015) using Personal Health Numbers. Concordant and discordant records were identified by whether or not a specific record co-existed in the two databases. Concordance and discordance (discrepancy) rates, i.e. percentage of concordant or discordant records, were estimated by years and drug types.
 ResultsDuring 2008-2015, there were 1,116,176 records collected by PIN, 1,005,548 records collected by ABC, and a total of 1,218,191 records collected by both for 13,413 ATP participants. The average discrepancy rate between PIN and ABC was 25.8%, and the rate was significantly lower for drugs commonly prescribed for health conditions in seniors, including cardiovascular diseases (18.7% for statin), hypertension (18.9% for beta blockers, ace inhibitors and diuretics), diabetes (23.8% for glucose lower drugs), COPD (20.2% for inhalers) and stomach disorders (22.3% for H2 antagonists and proton pump inhibitors), compared to other drugs (34.4%). For insured drugs, using ABC as reference database, 88.6% of ABC records were concordant with (co-existing in) PIN. The concordance rate for insured drug use was improved by 10% over 2008-2015.
 Conclusion/ImplicationsBy cross-linking two pharmaceutical databases in Alberta for senior ATP participants, we found remarkable discrepancies in pharmaceutical records between PIN and ABC, although there was noticeable improvement over the years. The discrepancy rate between PIN and ABC was drug-specific and significantly lower for drugs commonly prescribed in senior patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.003 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".