MétaCan
Menu
Back to cohort
Record W4385709848 · doi:10.1093/jamiaopen/ooad062

Automated identification of unstandardized medication data: a scalable and flexible data standardization pipeline using RxNorm on GEMINI multicenter hospital data

2023· article· en· W4385709848 on OpenAlexafffundabout
Riley Waters, Sarah Malecki, Sharan Lail, Denise Mak, Sudipta Saha, Hae Young Jung, Mohammed Arshad Imrit, Fahad Razak, Amol A. Verma

Bibliographic record

VenueJAMIA Open · 2023
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsUniversity of TorontoSt. Michael's Hospital
FundersAlliance de recherche numérique du Canada
KeywordsStandardizationComputer scienceIdentifierPharmacyCoding (social sciences)Data miningInformation retrievalMedicineStatisticsMathematicsFamily medicine

Abstract

fetched live from OpenAlex

Objective: Patient data repositories often assemble medication data from multiple sources, necessitating standardization prior to analysis. We implemented and evaluated a medication standardization procedure for use with a wide range of pharmacy data inputs across all drug categories, which supports research queries at multiple levels of granularity. Methods: The GEMINI-RxNorm system automates the use of multiple RxNorm tools in tandem with other datasets to identify drug concepts from pharmacy orders. GEMINI-RxNorm was used to process 2 090 155 pharmacy orders from 245 258 hospitalizations between 2010 and 2017 at 7 hospitals in Ontario, Canada. The GEMINI-RxNorm system matches drug-identifying information from pharmacy data (including free-text fields) to RxNorm concept identifiers. A user interface allows researchers to search for drug terms and returns the relevant original pharmacy data through the matched RxNorm concepts. Users can then manually validate the predicted matches and discard false positives. We designed the system to maximize recall (sensitivity) and enable excellent precision (positive predictive value) with efficient manual validation. We compared the performance of this system to manual coding (by a physician and pharmacist) of 13 medication classes. Results: Manual coding was performed for 1 948 817 pharmacy orders and GEMINI-RxNorm successfully returned 1 941 389 (99.6%) orders. Recall was greater than 0.985 in all 13 drug classes, and the F1-score and precision remained above 0.90 in all drug classes, facilitating efficient manual review to achieve 100% precision. GEMINI-RxNorm saved time substantially compared with manual standardization, reducing the time taken to review a pharmacy order row from an estimated 30 to 5 s and reducing the number of rows needed to be reviewed by up to 99.99%. Discussion and Conclusion: GEMINI-RxNorm presents a novel combination of RxNorm tools and other datasets to enable accurate, efficient, flexible, and scalable standardization of pharmacy data. By facilitating efficient manual validation, the GEMINI-RxNorm system can allow researchers to achieve near-perfect accuracy in medication data standardization.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.043
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.023
Threshold uncertainty score0.121

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.043
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0060.005
Science and technology studies0.0010.001
Scholarly communication0.0030.004
Open science0.0020.004
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.097
GPT teacher head0.405
Teacher spread0.308 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2023
Admission routes3
Has abstractyes

Explore more

Same venueJAMIA OpenSame topicBiomedical Text Mining and OntologiesFrench-language works237,207