MétaCan
Menu
Back to cohort
Record W4283364896 · doi:10.1101/2022.02.16.22268694

Automated identification of unstandardized medication data: A scalable and flexible data standardization pipeline using RxNorm on GEMINI multicenter hospital data

2022· preprint· en· W4283364896 on OpenAlexafffundabout
Riley Waters, Sarah Malecki, Sharan Lail, Denise Mak, Sudipta Saha, Hae Young Jung, Fahad Razak, Amol A. Verma

Bibliographic record

VenuemedRxiv · 2022
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicBiomedical Text Mining and Ontologies
Canadian institutionsUniversity of TorontoSt. Michael's Hospital
FundersNatural Sciences and Engineering Research Council of CanadaCanadian Institutes of Health ResearchAlliance de recherche numérique du CanadaCanadian Frailty NetworkUniversity of TorontoUniversity Health Network
KeywordsStandardizationComputer scienceIdentifierPharmacyFalse positive paradoxData miningCoding (social sciences)Information retrievalMedicineArtificial intelligenceStatisticsMathematicsFamily medicine

Abstract

fetched live from OpenAlex

ABSTRACT Objective Patient data repositories often assemble medication data from multiple sources, necessitating standardization prior to analysis. We implemented and evaluated a medication standardization procedure for use with a wide range of pharmacy data inputs across all drug categories, which supports research queries at multiple levels of granularity. Methods The GEMINI-RxNorm system automates the use of multiple RxNorm tools in tandem with other datasets to identify drug concepts from pharmacy orders. GEMINI-RxNorm was used to process 2,090,155 pharmacy orders from 245,258 hospitalizations between 2010 and 2017 at 7 hospitals in Ontario, Canada. The GEMINI-RxNorm system matches drug-identifying information from pharmacy data (including free-text fields) to RxNorm concept identifiers. A user interface allows researchers to search for drug terms and returns the relevant original pharmacy data through the matched RxNorm concepts. Users can then manually validate the predicted matches and discard false positives. We designed the system to maximize recall (sensitivity) and enable excellent precision (positive predictive value) with minimal manual validation. We compared the performance of this system to manual coding (by a physician and pharmacist) of 13 medication classes. Results Manual coding was performed for 1,948,817 pharmacy orders and GEMINI-RxNorm successfully returned 1,941,389 (99.6%) orders. Recall was greater than 98.5% in all 13 drug classes, and the F-Measure and precision remained above 90.0% in all drug classes, facilitating efficient manual review to achieve 100.0% precision. GEMINI-RxNorm saved time substantially compared to manual standardization, reducing the time taken to review a pharmacy order row from an estimated 30 seconds to 5 seconds and reducing the number of rows needed to be reviewed by up to 99.99%. Discussion and Conclusion GEMINI-RxNorm presents a novel combination of RxNorm tools and other datasets to enable accurate, efficient, flexible, and scalable standardization of pharmacy data. By facilitating efficient minimal manual validation, the GEMINI-RxNorm system can allow researchers to achieve near-perfect accuracy in medication data standardization.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.045
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.022
Threshold uncertainty score0.119

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.045
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0070.005
Science and technology studies0.0010.001
Scholarly communication0.0040.005
Open science0.0020.005
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.071
GPT teacher head0.370
Teacher spread0.299 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes3
Has abstractyes

Explore more

Same venuemedRxivSame topicBiomedical Text Mining and OntologiesFrench-language works237,207