Evaluating the Applicability and Appropriateness of ChatGPT as a Source for Tailored Nutrition Advice: A Multi-Scenario Study
Bibliographic record
Abstract
Background: In the rapidly evolving domain of healthcare technology, the integration of advanced computational models has opened up new possibilities for personalized nutrition guidance. The emergence of sophisticated language models, such as Chat Generative Pre-training Transformer (ChatGPT), offers potential in providing interactive and tailored dietary advice. However, concerns remain about the applicability and appropriateness of ChatGPT's recommendations, especially for those with distinct health conditions. Objectives: This study aimed to evaluate the reliability of ChatGPT as a source of nutritional advice. Methods: Three hypothetical scenarios representing various health conditions were presented alongside precise dietary requirements. ChatGPT was tasked to generate personalized dietary programs, encompassing meal timing, specific caloric portions (measured in grams and spoons), as well as alternative meal options for each scenario. Following this, ChatGPT’s generated dietary programs underwent a thorough review by a multidisciplinary team of nutritionist, specialist physicians and clinical researchers. The evaluation focused on the programs' suitability, alignment with dietary standards, consideration of individual health factors, and additional guidance Safety. Results: ChatGPT demonstrated its ability to generate various options of meal plans in accordance with basic nutrition principles. However, there are apparent issues with the recommended individual macronutrient distribution, handling health conditions, drug interactions, and setting realistic weight loss goals. Conclusions: While ChatGPT exhibits promise as a dietary program generator, its application for intervention should be restricted to certified nutrition professionals. Until July 2023, it is not advisable for patients to engage in self-prescription using ChatGPT version 3.5, owing to its inability to provide professional knowledge and acceptable guidance, particularly for individuals with co-existing conditions. The prevailing absence of clinical reasoning highlights the importance of employing ChatGPT solely as a tool, rather than relying on it as an autonomous decision-maker. Its lack of clinical reasoning highlighted the need for human intervention and expert collaboration for precise personalized evaluations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".