Linked Administrative Data at Statistics Canada – new data resources for horizontal research
Notice bibliographique
Résumé
There has been an increasing demand for analytics and research related to cross-cutting and horizontal issues in Canada, such as in the domains of housing, aging and immigration. Very often policy makers and stakeholders are posing a full spectrum of questions around a specific topic, requiring multidisciplinary evidence and data. Statistics Canada has a long history of record linkage. Over the past decade, the number of record linkage projects has increased exponentially. Several established platforms have been developed to facilitate linkage – Canadian Employer and Employer Database which brings together tax and employment records from both employees and employers; the Social Data Linkage Environment created to support linkages at the individuals level across a broad spectrum of social data (health, justice, education, socio-economic); and the Linkable File Environment for business data.
 The breadth of our data holdings married with record linkage capabilities allows the creation of data sets that crosses disciplines and areas or research. This presentation will showcase the innovative data integration approaches that Statistics Canada has advanced to meet the inter-disciplinary data needs.
 Statistics Canada are pioneering in some innovative linkages across various domains to help answer cross-cutting questions. For example, Longitudinal Administrative Databank linking longitudinal tax records to numerous other data files including tax records of spouses and children in the household, longitudinal Immigration Database linkage key and health records, is used to study economic impact of hospitalization, as well as better understand health outcomes of immigrants by various dimensions including socio-economic status. Other examples include the pilot projects linking Canadian Financial Capability Survey to tax records, to gauge the relationship between financial literacy and annual retirement savings behavior and Intergenerational Income Database being linked to Census to understand socio-economic factors affecting the intergenerational mobility.
 Rapid growth in data availability for research also poses new challenges on IM/IT, governance, access, capacity building, etc. As Statistics Canada has moved on a path of modernization, data integration is key to the development of new data sources to fill information gaps as we move forward.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,022 | 0,027 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,000 |
| Communication savante | 0,003 | 0,007 |
| Science ouverte | 0,027 | 0,011 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; les deux têtes enseignantes s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».