Validity of administrative health data case definitions for identifying polycystic ovary syndrome: a systematic review and meta-analysis
Bibliographic record
Abstract
STUDY QUESTION: What is the validity of published administrative health data case definitions of polycystic ovary syndrome (PCOS) compared with reference standards? SUMMARY ANSWER: Due to the limited number of eligible studies, drawing definitive conclusions is challenging; however, this review highlights significant gaps and variability in current PCOS case definitions, underscoring the need for standardized case definitions in future research. WHAT IS KNOWN ALREADY: Administrative health data offer the opportunity to evaluate health outcomes and disease epidemiology at a population-level. Currently, the validity of existing administrative health data case definitions for PCOS is unknown. STUDY DESIGN, SIZE, DURATION: A systematic review of the literature was conducted on full-text English-language articles up to July 2023, using the MEDLINE and EMBASE databases. PARTICIPANTS/MATERIALS, SETTING, METHODS: Two reviewers independently screened titles, abstracts and full texts, extracted data, assessed study quality and graded validity. A random effects meta-analysis was conducted to pool reported validity measures and heterogeneity was examined. MAIN RESULTS AND THE ROLE OF CHANCE: The review included four eligible articles consisting of three cross-sectional studies and one retrospective cohort study. Two studies defined PCOS using the Rotterdam Criteria, one study used self-report, and one used a clinical gold standard. All case definitions included the International Classification of Diseases (ICD)-9 code 256.4 for 'polycystic ovaries' and three studies used E28.2 for 'polycystic ovarian syndrome'. Three studies reported positive predictive value (PPV), which ranged from 30 to 96%. One study reported both PPV (96%) and sensitivity (50%) for one case definition. The pooled PPV estimate for the ICD code-based case definitions was 88% (95% confidence interval 82-95%; I2 = 100%). One study reported fair agreement (percent agreement= 90.3, κ = 0.27, percent agreement bias adjusted κ = 0.81). Overall, the risk of bias of the included studies was low. LIMITATIONS, REASONS FOR CAUTION: There were limited number of validations and precision indices of validations. WIDER IMPLICATIONS OF THE FINDINGS: Further validation of these case definitions in other administrative health datasets, and development of novel coding algorithms is required to inform future population-based studies in PCOS. STUDY FUNDING/COMPETING INTEREST(S): No external funding was used and there are no disclosures. REGISTRATION NUMBER: PROSPERO CRD42023385617.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.007 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".