Early Experience Analyzing Dietary Intake Data from the Canadian Community Health Survey—Nutrition Using the National Cancer Institute (NCI) Method
Bibliographic record
Abstract
BACKGROUND: One of the underpinning elements to support evidence-based decision-making in food and nutrition is the usual dietary intake of a population. It represents the long-run average consumption of a particular dietary component (i.e., food or nutrient). Variations in individual eating habits are observed from day-to-day and between individuals. The National Cancer Institute (NCI) method uses statistical modeling to account for these variations in estimation of usual intakes. This method was originally developed for nutrition survey data in the United States. The main objective of this study was to apply the NCI method in the analysis of Canadian nutrition surveys. METHODS: Data from two surveys, the 2004 and 2015 Canadian Community Health Survey-Nutrition were used to estimate usual dietary intake distributions from food sources using the NCI method. The effect of different statistical considerations such as choice of the model, covariates, stratification compared to pooling, and exclusion of outliers were assessed, along with the computational time to convergence. RESULTS: A flowchart to aid in model selection was developed. Different covariates (e.g., age/sex groups, cycle, weekday/weekend of the recall) were used to adjust the estimates of usual intakes. Moreover, larger differences in the ratio of within to between variation for a stratified analysis or a pooled analysis resulted in noticeable differences, particularly in the tails of the distribution of usual intake estimates. Outliers were subsequently removed when the ratio was larger than 10. For an individual age/sex group, the NCI method took 1 h-5 h to obtain results depending on the dietary component. CONCLUSION: Early experience in using the NCI method with Canadian nutrition surveys data led to the development of a flowchart to facilitate the choice of the NCI model to use. The ability of the NCI method to include covariates permits comparisons between both 2004 and 2015. This study shows that the improper application of pooling and stratification as well as the outlier detection can lead to biased results. This early experience can provide guidance to other researchers and ensures consistency in the analysis of usual dietary intake in the Canadian context.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.121 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.009 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.003 | 0.001 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".