Building partnerships, capacity, and knowledge through a use of newly linked child development and education datasets in Ontario, Canada.
Bibliographic record
Abstract
ObjectivesThe objective of this study was to establish a partnership between a university and a jurisdictional education body (Education Quality and Assessment Organization, EQAO) which would allow creation of a linked dataset from kindergarten to later grades in order to examine educational trajectory in mathematics in Ontario. ApproachBuilding on mutual goals of improving the understanding of children’s learning trajectories, we developed a project with an investigator team that included university researchers and representatives of the provincial educational assessment body, to link a database of child development status in kindergarten (Early Development Instrument/EDI data, including neighbourhood socioeconomic/SES index) with academic assessment EQAO data, and received research funding. A deterministic matching process was employed to match the datasets. We examined differences between the unmatched and fully matched cases and constructed a growth mixture model of math scores in grades 3, 6 and 9, with key EDI/SES variables as covariates. ResultsDespite lacking a common identifier, we successfully matched approximately 50% of the EDI cases from 2002-2014 (n=183,771). Effect sizes indicated negligible differences between matched and unmatched, except for SES and child development status, which were poorer for unmatched group. A 3-class solution was the best fit for a 20,000-person subsample of math trajectories based on AIC, BIC, ICL, and entropy values as well as sufficiently high proportions of posterior probabilities, which indicate confidence in class membership. 61% of sample showed steady moderate-high achievement; 9% started high, but declined, and 30% deteriorated then improved. Males, children in low SES, and those with adequate kindergarten EDI outcomes had better math achievement trajectories than females, children in high SES, and those with poor kindergarten outcomes. ConclusionGiven the two datasets were collected without explicit linkage plan, the matching was only 50%, nevertheless resulting in a large database that allows study of early development antecedents of students’ educational trajectories. The partnership between university and EQAO ensures a wide dissemination of results in both academia and policy worlds.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.044 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.007 | 0.015 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.002 | 0.006 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".