A Focused Review of Smartphone Diet-Tracking Apps: Usability, Functionality, Coherence With Behavior Change Theory, and Comparative Validity of Nutrient Intake and Energy Estimates
Bibliographic record
Abstract
: Smartphone diet-tracking apps may help individuals lose weight, manage chronic conditions, and understand dietary patterns; however, the usabilities and functionalities of these apps have not been well studied. : The aim of this study was to review the usability of current iPhone operating system (iOS) and Android diet-tracking apps, the degree to which app features align with behavior change constructs, and to assess variations between apps in nutrient coding. : The top 7 diet-tracking apps were identified from the iOS iTunes and Android Play online stores, downloaded and used over a 2-week period. Each app was independently scored by researchers using the System Usability Scale (SUS), and features were compared with the domains in an integrated behavior change theory framework: the Theoretical Domains Framework. An estimated 3-day food diary was completed using each app, and food items were entered into the United States Department of Agriculture (USDA) Food Composition Databases to evaluate their differences in nutrient data against the USDA reference. : Of the apps that were reviewed, LifeSum had the highest average SUS score of 89.2, whereas MyDietCoach had the lowest SUS score of 46.7. Some variations in features were noted between Android and iOS versions of the same apps, mainly for MyDietCoach, which affected the SUS score. App features varied considerably, yet all of the apps had features consistent with Beliefs about Capabilities and thus have the potential to promote self-efficacy by helping individuals track their diet and progress toward goals. None of the apps allowed for tracking of emotional factors that may be associated with diet patterns. The presence of behavior change domain features tended to be weakly correlated with greater usability, with R2 ranging from 0 to .396. The exception to this was features related to the Reinforcement domain, which were correlated with less usability. Comparing the apps with the USDA reference for a 3-day diet, the average differences were 1.4% for calories, 1.0% for carbohydrates, 10.4% for protein, and −6.5% for fat. : Almost all reviewed diet-tracking apps scored well with respect to usability, used a variety of behavior change constructs, and accurately coded calories and carbohydrates, allowing them to play a potential role in dietary intervention studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.011 | 0.059 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".