Measuring Caloric Intake at the Population Level (NOTION): Protocol for an Experimental Study
Bibliographic record
Abstract
BACKGROUND: The monitoring of caloric intake is an important challenge for the maintenance of individual and public health. The instruments used so far for dietary monitoring (eg, food frequency questionnaires, food diaries, and telephone interviews) are inexpensive and easy to implement but show important inaccuracies. Alternative methods based on wearable devices and wrist accelerometers have been proposed, yet they have limited accuracy in predicting caloric intake because analytics are usually not well suited to manage the massive sets of data generated from these types of devices. OBJECTIVE: This study aims to develop an algorithm using recent advances in machine learning methodology, which provides a precise and stable estimate of caloric intake. METHODS: The study will capture four individual eating activities outside the home over 2 months. Twenty healthy Italian adults will be recruited from the University of Padova in Padova, Italy, with email, flyers, and website announcements. The eligibility requirements include age 18 to 66 years and no eating disorder history. Each participant will be randomized to one of two menus to be eaten on weekdays in a predefined cafeteria in Padova (northeastern Italy). Flows of raw data will be accessed and downloaded from the wearable devices given to study participants and associated with anthropometric and demographic characteristics of the user (with their written permission). These massive data flows will provide a detailed picture of real-life conditions and will be analyzed through an up-to-date machine learning approach with the aim to accurately predict the caloric contribution of individual eating activities. Gold standard evaluation of the energy content of eaten foods will be obtained using calorimetric assessments made at the Laboratory of Dietetics and Nutraceutical Research of the University of Padova. RESULTS: The study will last 14 months from July 2017 with a final report by November 2018. Data collection will occur from October to December 2017. From this study, we expect to obtain a series of relevant data that, opportunely filtered, could allow the construction of a prototype algorithm able to estimate caloric intake through the recognition of food type and the number of bites. The algorithm should work in real time, be embedded in a wearable device, and able to match bite-related movements and the corresponding caloric intake with high accuracy. CONCLUSIONS: Building an automatic calculation method for caloric intake, independent on the black-box processing of the wearable devices marketed so far, has great potential both for clinical nutrition (eg, for assessing cardiovascular compliance or for the prevention of coronary heart disease through proper dietary control) and public health nutrition as a low-cost monitoring tool for eating habits of different segments of the population. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): DERR1-10.2196/12116.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.021 | 0.015 |
| Meta-epidemiology (narrow) | 0.006 | 0.003 |
| Meta-epidemiology (broad) | 0.007 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.005 | 0.003 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.006 | 0.007 |
| Insufficient payload (model declined to judge) | 0.072 | 0.021 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".