Development and validation of a simple algorithm to estimate common gestational age categories using standard administrative birth record data in Ontario, Canada
Bibliographic record
Abstract
Gestational age is often incompletely recorded in administrative records, despite being critical to paediatric and maternal health research. Several algorithms exist to estimate gestational age using administrative databases; however, many have not been validated or use complicated methods that are not readily adaptable. We developed a simple algorithm to estimate common gestational age categories from routine administrative data. We leveraged a population-based registry of all hospital births occurring in Ontario, Canada over 2002–2016 including 1.8 million birth records. In this sample, this simple algorithm had excellent performance compared to a verified measure of gestational age; 87.61% agreement (95% CI: 87.49, 87.74). The accuracy of the algorithm exceeded 98% for all of the gestational age categories. Agreement notably increased over time and was greatest among singleton births and infants born at 2500–2999 g. This study provides a straight-forward algorithm for accurately estimating common gestational age categories that is easily adaptable for use in other countries.Impact StatementWhat is already known on this subject? Gestational age is often incompletely or inaccurately recorded in administrative health databases, despite being critical to the study of many paediatric and maternal health outcomes. Consequently, researchers must rely on various methods to estimate gestational age, many of these methods are either overly simple (i.e. assuming a uniform duration) or analytically complicated and difficult to adapt for new populations (e.g. regression-based approaches).What the results of this study add? This study, based on a population-based registry of all 1.8 million births occurring in Ontario, Canada 2003–2016, found that a simple, sex-specific algorithm using three commonly recorded birth record characteristics performs almost perfectly compared to a clinical estimate recorded near birth.What the implications are of these findings for clinical practice and/or further research? This study suggests that a straight-forward, sex-specific algorithm based on routinely collected birth record data is able to accurately estimate common gestational age categories (i.e. extreme preterm, <28 weeks; very preterm, 28–32 weeks; moderate-to-late preterm, 33–26 weeks; and term, 37 weeks of completed gestational age). This work will be of greatest interest to perinatal researchers using routinely collected health administrative data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".