Linked Administrative Data at Statistics Canada – new data resources for horizontal research
Bibliographic record
Abstract
There has been an increasing demand for analytics and research related to cross-cutting and horizontal issues in Canada, such as in the domains of housing, aging and immigration. Very often policy makers and stakeholders are posing a full spectrum of questions around a specific topic, requiring multidisciplinary evidence and data. Statistics Canada has a long history of record linkage. Over the past decade, the number of record linkage projects has increased exponentially. Several established platforms have been developed to facilitate linkage – Canadian Employer and Employer Database which brings together tax and employment records from both employees and employers; the Social Data Linkage Environment created to support linkages at the individuals level across a broad spectrum of social data (health, justice, education, socio-economic); and the Linkable File Environment for business data. The breadth of our data holdings married with record linkage capabilities allows the creation of data sets that crosses disciplines and areas or research. This presentation will showcase the innovative data integration approaches that Statistics Canada has advanced to meet the inter-disciplinary data needs. Statistics Canada are pioneering in some innovative linkages across various domains to help answer cross-cutting questions. For example, Longitudinal Administrative Databank linking longitudinal tax records to numerous other data files including tax records of spouses and children in the household, longitudinal Immigration Database linkage key and health records, is used to study economic impact of hospitalization, as well as better understand health outcomes of immigrants by various dimensions including socio-economic status. Other examples include the pilot projects linking Canadian Financial Capability Survey to tax records, to gauge the relationship between financial literacy and annual retirement savings behavior and Intergenerational Income Database being linked to Census to understand socio-economic factors affecting the intergenerational mobility. Rapid growth in data availability for research also poses new challenges on IM/IT, governance, access, capacity building, etc. As Statistics Canada has moved on a path of modernization, data integration is key to the development of new data sources to fill information gaps as we move forward.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.033 | 0.156 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.027 | 0.069 |
| Science and technology studies | 0.007 | 0.002 |
| Scholarly communication | 0.015 | 0.007 |
| Open science | 0.007 | 0.011 |
| Research integrity | 0.003 | 0.006 |
| Insufficient payload (model declined to judge) | 0.060 | 0.029 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".