Linked data and inclusion health: Harmonised international data linkage to identify determinants of health inequalities
Bibliographic record
Abstract
A recent article in The Lancet establishing the principles of inclusion health, highlighted substantial gaps in our understanding of the drivers of health inequalities in socially excluded groups such as people with a history of incarceration, people who experience homelessness, sex workers, people with mental illness, and people who inject drugs1. Cross-sectoral data linkage of electronic health records with services working with socially excluded groups was one of the key recommendations of this article. The magnitude of health disparities observed in people that experience social exclusion necessitates an international public health response and addressing the determinants of social exclusion has been identified as a key component of closing the gap of Indigenous disadvantage2. This symposium will establish data linkage as a key component of the inclusion health and will complement the efforts of the Pan American Health Oranization's (PAHO) Commission on Equity and Health Inequalities in the Americas. Traditional survey methodology is costly and often results in studies that are highly parochial in nature. Due to difficulties recruiting and retaining marginalized groups, these studies are commonly forced to adopt methodological concessions, often selecting the most convenient participants (i.e., selection bias) or incurring increased rates of loss-to-follow-up (i.e., attrition bias). Conversely, global studies aimed at modelling the burden of disease are often not sufficiently nuanced to answer specific inferential research questions. Data-linkage has the potential to overcome these common biases and limitations. Thus, harmonised international data-linkage studies are an important component of the inclusion health response to identify the determinants of health inequalities in socially excluded groups and inform the global inclusion health agenda. This symposium will bring together facilitators from three countries with extensive experience conducting data linkage studies that generate evidence on health and social inequality in socially excluded groups. Using a current multinational study as an example, barriers to international data-linkage studies, methodological solutions, and distributed approaches to generating international comparative evidence will be presented. Innovative examples of cross-sectoral approaches to linkage with social service, correctional and national survey data will be discussed. The development of a novel framework for identifying social exclusion exposures and determinants of health inequalities typically not captured in administrative health data will also be discussed. The session will conclude with a discussion aimed at forming the foundation of an international data linkage project to address these current gaps identified in the inclusion health series and best practice for translation to policy and practice to address health disparities in socially excluded groups. References Aldridge et al. Morbidity and mortality in homeless individuals, prisoners, sex workers, and individuals with substance use disorders in high-income countries: a systematic review and meta-analysis. The Lancet. 2017;391(10117):241-250. https://doi.org/10.1016/S0140-6736(17)31869-X Greenwood M et al. Challenges in health equity for Indigenous peoples in Canada. The Lancet. 2018;Epub ahead of print. https://doi.org/10.1016/S0140-6736(18)30177-6
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.405 | 0.528 |
| Meta-epidemiology (narrow) | 0.003 | 0.003 |
| Meta-epidemiology (broad) | 0.006 | 0.007 |
| Bibliometrics | 0.022 | 0.036 |
| Science and technology studies | 0.005 | 0.006 |
| Scholarly communication | 0.016 | 0.016 |
| Open science | 0.009 | 0.060 |
| Research integrity | 0.007 | 0.011 |
| Insufficient payload (model declined to judge) | 0.016 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".