The Utility of Different Data Standards to Document Adverse Drug Event Symptoms and Diagnoses: Mixed Methods Study
Bibliographic record
Abstract
BACKGROUND: Existing systems to document adverse drug events often use free text data entry, which produces nonstandardized and unstructured data that are prone to misinterpretation. Standardized terminology may improve data quality; however, it is unclear which data standard is most appropriate for documenting adverse drug event symptoms and diagnoses. OBJECTIVE: This study aims to compare the utility, strengths, and weaknesses of different data standards for documenting adverse drug event symptoms and diagnoses. METHODS: We performed a mixed methods substudy of a multicenter retrospective chart review. We reviewed the research records of prospectively diagnosed adverse drug events at 5 Canadian hospitals. A total of 2 pharmacy research assistants independently entered the symptoms and diagnoses for the adverse drug events using four standards: Medical Dictionary for Regulatory Activities (MedDRA), Systematized Nomenclature of Medicine (SNOMED) Clinical Terms, SNOMED Adverse Reaction (SNOMED ADR), and International Classification of Diseases (ICD) 11th Revision. Disagreements between research assistants regarding the case-specific utility of data standards were discussed until a consensus was reached. We used consensus ratings to determine the proportion of adverse drug events covered by a data standard and coded and analyzed field notes from the consensus sessions. RESULTS: We reviewed 573 adverse drug events and found that MedDRA and ICD-11 had excellent coverage of adverse drug event symptoms and diagnoses. MedDRA had the highest number of matches between the research assistants, whereas ICD-11 had the fewest. SNOMED ADR had the lowest proportion of adverse drug event coverage. The research assistants were most likely to encounter terminological challenges with SNOMED ADR and usability challenges with ICD-11, whereas least likely to encounter challenges with MedDRA. CONCLUSIONS: Usability, comprehensiveness, and accuracy are important features of data standards for documenting adverse drug event symptoms and diagnoses. On the basis of our results, we recommend the use of MedDRA.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.227 | 0.313 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.006 |
| Bibliometrics | 0.010 | 0.007 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".