Linkage of whole genome sequencing and administrative health data in autism: A proof of concept study
Bibliographic record
Abstract
Whether genetic testing in autism can help understand longitudinal health outcomes and health service needs is unclear. The objective of this study was to determine whether carrying an autism-associated rare genetic variant is associated with differences in health system utilization by autistic children and youth. This retrospective cohort study examined 415 autistic children/youth who underwent genome sequencing and data collection through a translational neuroscience program (Province of Ontario Neurodevelopmental Disorders Network). Participant data were linked to provincial health administrative databases to identify historical health service utilization, health care costs, and complex chronic medical conditions during a 3-year period. Health administrative data were compared between participants with and without a rare genetic variant in at least 1 of 74 genes associated with autism. Participants with a rare variant impacting an autism-associated gene (n = 83, 20%) were less likely to have received psychiatric care (at least one psychiatrist visit: 19.3% vs. 34.3%, p = 0.01; outpatient mental health visit: 66% vs. 77%, p = 0.04). Health care costs were similar between groups (median: $5589 vs. $4938, p = 0.4) and genetic status was not associated with odds of being a high-cost participant (top 20%) in this cohort. There were no differences in the proportion with complex chronic medical conditions between those with and without an autism-associated genetic variant. Our study highlights the feasibility and potential value of genomic and health system data linkage to understand health service needs, disparities, and health trajectories in individuals with neurodevelopmental conditions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.059 | 0.120 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.003 | 0.002 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".