P770 Ascertainment of pediatric inflammatory bowel disease cases from administrative health data in Québec, Canada
Bibliographic record
Abstract
Abstract Background Administrative databases are useful for estimating population-level disease occurrence. Our objective was to ascertain cases of pediatric inflammatory bowel disease (IBD) by applying two validated algorithms to administrative health data, evaluate agreement, and compare health services utilisation between concordant and discordant cases. Methods The Quebec Birth Cohort on Immunity Health was established through linkage of administrative databases and includes 400 611 persons born in the province of Québec (Canada) from 1970 to 1974. Physician consultations (PC) and hospitalisations (H) for IBD were documented in health databases until 2014. Two validated algorithms were used to identify pediatric IBD cases. Firstly, a single-step algorithm was applied [5PC or 2H within 4 years]. Secondly, a two-step algorithm was implemented, first considering whether the person had undergone sigmoidoscopy/colonoscopy before age 18 [yes: 4PC or 2H within 3 years; no: 7PC or 3H within 3 years]. We evaluated the agreement between both algorithms using the Kappa statistic, and compared health services utilisation among concordant and discordant cases using a t-test. Results The single-step algorithm generated 527 pediatric IBD cases (0.13%), whereas 480 (0.12%) were identified with the multi-step algorithm. Among the 534 cases identified by either algorithm, 473 (88.6%) were identified by both, 54 (10.1%) only by the single-step, and 7 (1.3%) only by the multi-step algorithm. Kappa was 0.94 (95% confidence interval: 0.92, 0.95), and the proportions of positive and negative agreement were respectively 0.94 and 1.00. The average number of PC and H before age 18 years among concordant and discordant cases was respectively 26.0 and 3.9 (p < 0.0001). Conclusion The prevalence of pediatric IBD was similar when applying two different case identification algorithms, few cases were discordant. In the near future, a survey conducted in a subset of the cohort will allow us to compare self-report with ascertainment from administrative databases.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.007 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".