Validation of Insurance Billing Codes for Monitoring Antenatal Screening
Bibliographic record
Abstract
BACKGROUND: Prevalence statistics for pregnancy complications identified through screening such as gestational diabetes usually assume universal screening. However, rates of screening completion in pregnancy are not available in many birth registries or hospital databases. We validated screening-test completion by comparing public insurance laboratory and radiology billing records with medical records at three hospitals in British Columbia, Canada. METHODS: We abstracted a random sample of 140 delivery medical records (2014-2019), and successfully linked 127 to valid provincial insurance billings and maternal-newborn registry data. We compared billing records for gestational diabetes screening, any ultrasound before 14 weeks gestational age, and Group B streptococcus screening during each pregnancy to the gold standard of medical records by calculating sensitivity and specificity, positive predictive value, negative predictive value, and prevalence with 95% confidence intervals (CIs). RESULTS: Gestational diabetes screening (screened vs. unscreened) in billing records had a high sensitivity (98% [95% CI = 93, 100]) and specificity (>99% [95% CI = 86, 100]). The use of specific glucose screening approaches (two-step vs. one-step) were also well characterized by billing data. Other tests showed high sensitivity (ultrasound 97% [95% CI = 92, 99]; Group B streptococcus 96% [95% CI = 89, 99]) but lower negative predictive values (ultrasound 64% [95% CI = 33, 99]; Group B streptococcus 70% [95% CI = 40, 89]). Lower negative predictive values were due to the high prevalence of these screening tests in our sample. CONCLUSIONS: Laboratory and radiology insurance billing codes accurately identified those who completed routine antenatal screening tests with relatively low false-positive rates.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.059 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".