Validation of Insurance Billing Codes for Monitoring Antenatal Screening
Bibliographic record
Abstract
BACKGROUND: Prevalence statistics for pregnancy complications identified through screening such as gestational diabetes usually assume universal screening. However, rates of screening completion in pregnancy are not available in many birth registries or hospital databases. We validated screening-test completion by comparing public insurance laboratory and radiology billing records with medical records at three hospitals in British Columbia, Canada. METHODS: We abstracted a random sample of 140 delivery medical records (2014-2019), and successfully linked 127 to valid provincial insurance billings and maternal-newborn registry data. We compared billing records for gestational diabetes screening, any ultrasound before 14 weeks gestational age, and Group B streptococcus screening during each pregnancy to the gold standard of medical records by calculating sensitivity and specificity, positive predictive value, negative predictive value, and prevalence with 95% confidence intervals (CIs). RESULTS: Gestational diabetes screening (screened vs. unscreened) in billing records had a high sensitivity (98% [95% CI = 93, 100]) and specificity (>99% [95% CI = 86, 100]). The use of specific glucose screening approaches (two-step vs. one-step) were also well characterized by billing data. Other tests showed high sensitivity (ultrasound 97% [95% CI = 92, 99]; Group B streptococcus 96% [95% CI = 89, 99]) but lower negative predictive values (ultrasound 64% [95% CI = 33, 99]; Group B streptococcus 70% [95% CI = 40, 89]). Lower negative predictive values were due to the high prevalence of these screening tests in our sample. CONCLUSIONS: Laboratory and radiology insurance billing codes accurately identified those who completed routine antenatal screening tests with relatively low false-positive rates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".