Multilevel Estimation of the Relative Impacts of Social Determinants on Income-Related Health Inequalities in Urban Canada: Protocol for the Canadian Social Determinants Urban Laboratory (Preprint)
Bibliographic record
Abstract
BACKGROUND Two decades of research have highlighted persistent income-related health inequities in Canada across municipal, provincial, and national levels. While there is broad consensus among researchers, advocates, and health professionals that social determinants are the primary drivers of health, the empirical foundation supporting this remains relatively limited. A current renaissance in health system data access offers an opportunity to assess the multilevel impact of social factors on health inequalities, yet this potential remains underused. OBJECTIVE This project aims to examine how social, economic, and political conditions shape health inequalities and investigate how structural and intermediate determinants explain disparities across national, provincial, city, neighborhood, and individual levels. METHODS We will create the Canadian Social Determinants Urban Laboratory (CSDUL), a multilevel, longitudinal, virtual data environment that integrates 15 existing databases from Statistics Canada, the Canadian Institute for Health Information, the Canadian Urban Environmental Health Research Consortium, and DMTI Spatial. Guided by the World Health Organization social determinants of health framework, CSDUL will initially cover 2011 to 2015 due to data completeness and expand as additional years become available. CSDUL builds on Statistics Canada’s Canadian Population Health Survey and will link survey data to administrative and health records, including hospital discharges, ambulatory care, mortality, cancer registries, and longitudinal tax files. Area-level indicators will be added using historical postal codes and geospatial boundaries. Organized through a hub-and-node model, CSDUL includes a central hub and 5 research nodes. We will develop and validate area-based indicators to study social determinants at micro (individual), meso (neighborhood, city, and province), and macro (national) levels. A core deliverable is to assess the strengths and limitations of survey and administrative data for health research and derive variables accordingly. After developing CSDUL, we will replicate World Health Organization Regional Office for Europe income-related health inequality analysis for urban Canada and analyze the impact of social determinants on outcomes. We will apply a 2-fold Oaxaca-Blinder decomposition between the lowest and highest urban income quintiles. A major strength of CSDUL is its capacity to analyze how diverse determinants shape health across subgroups (eg, gender), identifying key drivers of health outcomes. RESULTS The indicators to be used in CSDUL are being developed and validated by the contributing nodes. In collaboration with node 3, we are constructing measures of social capital using DMTI Spatial Points of Interest data. A prototype version of CSDUL incorporating a limited set of indicators has been developed in Statistics Canada’s Research Data Centre. We anticipate receiving the finalized indicators from the nodes by August 2025 to September 2025 and aim to complete the decomposition analysis by December 2025. CONCLUSIONS Multisectoral interventions are most effective when they are customized to meet the unique needs of specific subpopulations using robust and multilevel data sources such as CSDUL. INTERNATIONAL REGISTERED REPORT DERR1-10.2196/71929
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.060 |
| Meta-epidemiology (narrow) | 0.002 | 0.003 |
| Meta-epidemiology (broad) | 0.002 | 0.004 |
| Bibliometrics | 0.004 | 0.007 |
| Science and technology studies | 0.009 | 0.002 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.116 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".