NutritionVerse3D2D: Large 3D Object and 2D Image Food Dataset for Dietary Intake Estimation
Bibliographic record
Abstract
Elderly populations often face significant challenges when it comes to dietary intake tracking, often exacerbated by health complications. Unfortunately, conventional diet assessment techniques such as food frequency questionnaires, food diaries, and 24 h recall are subject to substantial bias. Recent advancements in machine learning and computer vision show promise of automated nutrition tracking methods of food, but require a large, high-quality dataset in order to accurately identify the nutrients from the food on the plate. However, manual creation of large-scale datasets with such diversity is time-consuming and hard to scale. On the other hand, synthesized 3D food models enable view augmentation to generate countless photorealistic 2D renderings from any viewpoint, reducing imbalance across camera angles. In this paper, we present a process to collect a large image dataset of food scenes that span diverse viewpoints and highlight its usage in dietary intake estimation. We first collect quality 3D objects of food items (NV-3D) that are used to generate photorealistic synthetic 2D food images (NV-Synth) and then manually collect a validation 2D food image dataset (NV-Real). We benchmark various intake estimation approaches on these datasets and present NutritionVerse3D2D, a collection of datasets that contain 3D objects and 2D images, along with models that estimate intake from the 2D food images. We release all the datasets along with the developed models to accelerate machine learning research on dietary sensing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.002 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".