Systematic Review and Critical Appraisal of Validation Studies to Identify Rheumatic Diseases in Health Administrative Databases
Bibliographic record
Abstract
OBJECTIVE: To evaluate the quality of the methods and reporting of published studies that validate administrative database algorithms for rheumatic disease case ascertainment. METHODS: We systematically searched MEDLINE, Embase, and the reference lists of articles published from 1980 to 2011. We included studies that validated administrative data algorithms for rheumatic disease case ascertainment using medical record or patient-reported diagnoses as the reference standard. Each study was evaluated using published standards for the reporting and quality assessment of diagnostic accuracy, which informed the development of a methodologic framework to help critically appraise and guide research in this area. RESULTS: Twenty-three studies met the inclusion criteria. Administrative database algorithms to identify cases were most frequently validated against diagnoses in medical records (83%). Almost two-thirds of the studies (61%) used diagnosis codes in administrative data to identify potential cases and then reviewed medical records to confirm the diagnoses. The remaining studies did the reverse, identifying patients using a reference standard and then testing algorithms to identify cases in administrative data. Many authors (61%) described the patient population, but few (26%) reported key measures of diagnostic accuracy (sensitivity, specificity, and positive and negative predictive values). Only one-third of studies reported disease prevalence in the validation study sample. CONCLUSION: The methods used in administrative data validation studies of rheumatic diseases are highly variable. Few studies reported key measures of diagnostic accuracy despite their importance for drawing conclusions about the validity of administrative database algorithms. We developed a methodologic framework and recommendations for validation study conduct and reporting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.288 | 0.655 |
| Meta-epidemiology (narrow) | 0.005 | 0.004 |
| Meta-epidemiology (broad) | 0.021 | 0.015 |
| Bibliometrics | 0.044 | 0.026 |
| Science and technology studies | 0.003 | 0.005 |
| Scholarly communication | 0.008 | 0.007 |
| Open science | 0.008 | 0.005 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".