Systematic Review and Critical Appraisal of Validation Studies to Identify Rheumatic Diseases in Health Administrative Databases
Bibliographic record
Abstract
OBJECTIVE: To evaluate the quality of the methods and reporting of published studies that validate administrative database algorithms for rheumatic disease case ascertainment. METHODS: We systematically searched MEDLINE, Embase, and the reference lists of articles published from 1980 to 2011. We included studies that validated administrative data algorithms for rheumatic disease case ascertainment using medical record or patient-reported diagnoses as the reference standard. Each study was evaluated using published standards for the reporting and quality assessment of diagnostic accuracy, which informed the development of a methodologic framework to help critically appraise and guide research in this area. RESULTS: Twenty-three studies met the inclusion criteria. Administrative database algorithms to identify cases were most frequently validated against diagnoses in medical records (83%). Almost two-thirds of the studies (61%) used diagnosis codes in administrative data to identify potential cases and then reviewed medical records to confirm the diagnoses. The remaining studies did the reverse, identifying patients using a reference standard and then testing algorithms to identify cases in administrative data. Many authors (61%) described the patient population, but few (26%) reported key measures of diagnostic accuracy (sensitivity, specificity, and positive and negative predictive values). Only one-third of studies reported disease prevalence in the validation study sample. CONCLUSION: The methods used in administrative data validation studies of rheumatic diseases are highly variable. Few studies reported key measures of diagnostic accuracy despite their importance for drawing conclusions about the validity of administrative database algorithms. We developed a methodologic framework and recommendations for validation study conduct and reporting.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.028 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".