The Generalizability of a Medication Administration Discrepancy Detection System: Quantitative Comparative Analysis
Bibliographic record
Abstract
BACKGROUND: As a result of the overwhelming proportion of medication errors occurring each year, there has been an increased focus on developing medication error prevention strategies. Recent advances in electronic health record (EHR) technologies allow institutions the opportunity to identify medication administration error events in real time through computerized algorithms. MED.Safe, a software package comprising medication discrepancy detection algorithms, was developed to meet this need by performing an automated comparison of medication orders to medication administration records (MARs). In order to demonstrate generalizability in other care settings, software such as this must be tested and validated in settings distinct from the development site. OBJECTIVE: The purpose of this study is to determine the portability and generalizability of the MED.Safe software at a second site by assessing the performance and fit of the algorithms through comparison of discrepancy rates and other metrics across institutions. METHODS: The MED.Safe software package was executed on medication use data from the implementation site to generate prescribing ratios and discrepancy rates. A retrospective analysis of medication prescribing and documentation patterns was then performed on the results and compared to those from the development site to determine the algorithmic performance and fit. Variance in performance from the development site was further explored and characterized. RESULTS: Compared to the development site, the implementation site had lower audit/order ratios and higher MAR/(order + audit) ratios. The discrepancy rates on the implementation site were consistently higher than those from the development site. Three drivers for the higher discrepancy rates were alternative clinical workflow using orders with dosing ranges; a data extract, transfer, and load issue causing modified order data to overwrite original order values in the EHRs; and delayed EHR documentation of verbal orders. Opportunities for improvement were identified and applied using a software update, which decreased false-positive discrepancies and improved overall fit. CONCLUSIONS: The execution of MED.Safe at a second site was feasible and effective in the detection of medication administration discrepancies. A comparison of medication ordering, administration, and discrepancy rates identified areas where MED.Safe could be improved through customization. One modification of MED.Safe through deployment of a software update improved the overall algorithmic fit at the implementation site. More flexible customizations to accommodate different clinical practice patterns could improve MED.Safe's fit at new sites.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.170 | 0.444 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.005 | 0.003 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".