Clustering of Health Behaviors in Canadians: A Multiple Behavior Analysis of Data from the Canadian Longitudinal Study on Aging
Bibliographic record
Abstract
BACKGROUND: Health behaviors such as physical inactivity, unhealthy eating, smoking tobacco, and alcohol use are each leading risk factors for non-communicable chronic disease. Better understanding which behaviors tend to co-occur (i.e., cluster together) and co-vary (i.e., are correlated) may provide novel opportunities to develop more comprehensive interventions to promote multiple health behavior change. However, whether co-occurrence or co-variation-based approaches are better suited for this task remains relatively unknown. PURPOSE: To compare the utility of co-occurrence vs. co-variation-based approaches for understanding the interconnectedness between multiple health-impacting behaviors. METHODS: Using baseline and follow-up data (N = 40,268) from the Canadian Longitudinal Study of Aging, we examined the co-occurrence and co-variation of health behaviors. We used cluster analysis to group individuals based on their behavioral tendencies across multiple behaviors and to examine how these clusters are associated with demographic characteristics and health indicators. We compared outputs from cluster analysis to behavioral correlations and compared regression analyses of clusters and individual behaviors predicting future health outcomes. RESULTS: Seven clusters were identified, with clusters differentiated by six of the seven health behaviors included in the analysis. Sociodemographic characteristics varied across several clusters. Correlations between behaviors were generally small. In regression analyses individual behaviors accounted for more variance in health outcomes than clusters. CONCLUSIONS: Co-occurrence-based approaches may be more suitable for identifying sub-groups for intervention targeting while co-variation approaches are more suitable for building an understanding of the relationships between health behaviors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".