Data accuracy in the Ontario birth Registry: a chart re-abstraction study
Bibliographic record
Abstract
BACKGROUND: Ontario's birth Registry (BORN) was established in 2009 to collect, interpret, and share critical data about pregnancy, birth and the early childhood period to facilitate and improve the provision of healthcare. Since the use of routinely-collected health data has been prioritized internationally by governments and funding agencies to improve patient care, support health system planning, and facilitate epidemiological surveillance and research, high quality data is essential. The purpose of this study was to verify the accuracy of a selection of data elements that are entered in the Registry. METHODS: Data quality was assessed by comparing data re-abstracted from patient records to data entered into the Ontario birth Registry. A purposive sample of 10 hospitals representative of hospitals in Ontario based on level of care, birth volume and geography was selected and a random sample of 100 linked mother and newborn charts were audited for each site. Data for 29 data elements were compared to the corresponding data entered in the Ontario birth Registry using percent agreement, kappa statistics for categorical data elements and intra-class correlation coefficients (ICCs) for continuous data elements. RESULTS: Agreement ranged from 56.9 to 99.8%, but 76% of the data elements (22 of 29) had greater than 90% agreement. There was almost perfect (kappa 0.81-0.99) or substantial (kappa 0.61-0.80) agreement for 12 of the categorical elements. Six elements showed fair-to-moderate agreement (kappa <0.60). We found moderate-to-excellent agreement for four continuous data elements (ICC >0.50). CONCLUSION: Overall, the data elements we evaluated in the birth Registry were found to have good agreement with data from the patients' charts. Data elements that showed moderate kappa or low ICC require further investigation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.048 | 0.182 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.012 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".