Associations between Dietary Pattern Networks Derived from Machine Learning Algorithms and Cardiovascular Disease Risk in the NutriNet-Santé Cohort
Bibliographic record
Abstract
BACKGROUND: Major advances in the fields of data science and machine learning have enabled the use of novel methods, such as Gaussian graphical models (GGMs) and the Louvain algorithm, to identify dietary patterns (DP). OBJECTIVES: The aim of this study was to identify DP networks using novel computational approaches and to investigate the associations between these DP networks and cardiovascular disease (CVD) risk in a sample of the French population. METHODS: A sample of 99,362 participants aged ≥15 y from the NutriNet-Santé cohort was used. Dietary intakes (reported as grams per day) were assessed using ≥2 24-h dietary records, which were then classified into 42 food groups. CVD events were assessed using health questionnaires and subsequently validated based on medical records. GGMs were employed with the Louvain algorithm to derive DP networks. GGMs are network models that depict relationships among many variables (food groups) based on conditional correlation matrices. The Louvain algorithm extracts nonoverlapping communities from large networks. The relationship between DP networks and CVD incidence was evaluated using proportional hazard Cox models, adjusted for confounding variables. RESULTS: Analyses revealed 5 distinct DP networks reflecting consumption of 1) appetizer foods, 2) breakfast foods, 3) plant-based foods, 4) ultraprocessed sweets and snacks, and 5) healthy foods. Among these, only the DP network of ultraprocessed sweets and snacks was associated with greater CVD risk when adjusted for energy and potential confounders including overall diet quality (hazard ratio of quintile 5 compared with quintile 1: 1.32; 95% confidence interval: 1.11, 1.57; P-trend = 0.0002). CONCLUSIONS: The results suggest that a DP network reflecting the consumption of ultraprocessed sweets and snacks is associated with incident CVD in a sample of the French population, independent of diet quality. The innovative approach to derive empirical DP networks may assist in the identification of food groups that are likely to be consumed together in a population, thereby helping to identify dietary habits to target for the prevention of CVD. This trial was registered at clinicaltrials.gov as NCT03335644.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".