Multinational Population-Based Health Surveys Linked to Outcome Data: An Untapped Resource
Bibliographic record
Abstract
ABSTRACTObjectivesIncreasingly, national health surveys are being linked to vital statistics and health care information, providing a new and unique source of individual population health data. Given that nationally-representative health surveys are performed in over a hundred countries, these linkages create comprehensive data sets that are potentially larger than most existing cohort studies. To date, this resource has not been utilized.
 ApproachThe purpose, study base, content and methods of the Canadian Community Health Survey (CCHS) cycles 2.1 (2003-04) and 3.1 (2005-06) and the United States 2000 and 2005 National Health Interview Survey (NHIS) were examined for comparability and consistency. Smoking, alcohol, physical activity and diet questions were identified, question construct and response categorization were compared, and variable constructions possible for both national health surveys were created. All respondents 20+ years of age were identified and stratified by country and sex. Cox proportional hazard models were used to estimate 5-year hazards of mortality associated with the common smoking, alcohol, physical activity and diet variables.
 ResultsThe CCHS and NHIS are highly consistent and comparable. Health behaviour questions are similar and permit the creation of smoking, alcohol, physical activity and diet variables that are comparable across surveys. A total of 284 475 survey respondents from Canada and the United States (CCHS, N= 226 731; NHIS, N= 57 744) were included. The largest mortality hazards were associated with female heavy smokers in both Canada (HR: 2.91; 95% CI: 2.52, 3.37) and the United States (Female HR: 2.96; 95% CI: 2.59, 3.38), compared to non-smokers. Moderate variation in the age adjusted all-cause mortality hazard ratios was observed; both smoking and physical activity hazard ratios were consistently higher in the United States than in Canada.
 ConclusionThis study provides initial support for the methodological feasibility of pooling linked population health surveys however, challenges introduced by dissimilarities will require the use of innovative methodologies, and discussions regarding how to manage jurisdictional data restrictions and privacy issues are needed. Pooled population health data has the potential to improve national and international health surveillance and public health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.022 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.001 | 0.004 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".