Systematic Review of Validation Studies of the Use of Administrative Data to Identify Serious Infections
Bibliographic record
Abstract
OBJECTIVE: To conduct a systematic review of the literature on the validation of algorithms identifying infections in administrative data for future use in populations with rheumatic diseases. METHODS: Medline and EMBase were searched using the themes "administrative data" and "infection" between 1950 and October 2012. Inclusion criteria consisted of validation studies of administrative data identifying infections in adult populations. Article quality was assessed using a validated tool. RESULTS: A total of 5,941 articles were identified, 90 articles underwent detailed review, and 24 studies were included. The majority (17 of 24) examined bacterial infections and 9 examined opportunistic infections. Eighteen studies were from the US and all but 4 studies used International Classification of Diseases, Ninth Revision codes. Rheumatoid arthritis patients were studied in 6 of 24 articles. The studies on bacterial infections in general reported highly variable sensitivity and positive predictive value (PPV) for the diagnosis of infections using administrative data (sensitivity range 4.4-100%, PPV range 21.7-100%). Algorithms to identify opportunistic infections similarly had a highly variable sensitivity (range 20-100%) and PPV (range 1.3-100%). Thirteen studies compared the diagnostic accuracy of different algorithms, which revealed that strategies including a comprehensive algorithm using a greater number of diagnostic codes or codes in any position had the highest sensitivity for the diagnosis of infection. Algorithms that incorporated microbiologic or pharmacy data in combination with diagnostic codes had improved PPV for identification of tuberculosis. CONCLUSION: Algorithms for identifying infections using administrative data should be selected based on the purpose of the study, with careful consideration as to whether a high sensitivity or PPV is required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".