Validating a popular outpatient antibiotic database to reliably identify high prescribing physicians for patients 65 years of age and older
Bibliographic record
Abstract
OBJECTIVE: Many jurisdictions lack comprehensive population-based antibiotic use data and rely on third party companies, most commonly IQVIA. Our objective was to validate the accuracy of the IQVIA Xponent antibiotic database in identifying high prescribing physicians compared to the reference standard of a highly accurate population-wide database of outpatient antimicrobial dispensing for patients ≥65 years. METHODS: We conducted this study between 1 March 2016 and 28 February 2017 in Ontario, Canada. We evaluated the agreement and correlation between the databases using kappa statistics and Bland-Altman plots. We also assessed performance characteristics for Xponent to accurately identify high prescribing physicians with sensitivity, specificity, positive predictive value (PPV), and negative predictive value. RESULTS: We included 9,272 physicians. The Xponent database has a specificity of 92.4% (95%CI 92.0%-92.8%) and PPV of 77.2% (95%CI 76.0%-78.4%) for correctly identifying the top 25th percentile of physicians by antibiotic volume. In the sensitivity analysis, 94% of the top 25th percentile physicians in Xponent were within the top 40th percentile in the reference database. The mean number of antibiotic prescriptions per physician were similar with a relative difference of -0.4% and 2.7% for female and male patients, respectively. The error was greater in rural areas with a relative difference of -8.4% and -5.6% per physician for female and male patients, respectively. The weighted kappa for quartile agreement was 0.68 (95%CI 0.67-0.69). CONCLUSION: We validated the IQVIA Xponent antibiotic database to identify high prescribing physicians for patients ≥65 years, and identified some important limitations. Collecting accurate population-based antibiotic use data will remain vital to global antimicrobial stewardship efforts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".