MétaCan
Menu
Back to cohort
Record W4408612519 · doi:10.2196/60528

Natural Language Processing and Machine Learning Techniques for Analyzing Conversations About Nutritional Yeasts in the United States and France: Retrospective Social Media Listening Study

2025· article· en· W4408612519 on OpenAlexvenueno aff
Jean-François Jeanne, Joelle Malaab, Antoine Vanhove, Florian Mourey, Manissa Talmatkadi, Stéphane Schück

Bibliographic record

VenueJMIR Infodemiology · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicSocial Media in Health Education
Canadian institutionsnot available
Fundersnot available
KeywordsPreprintComputer scienceNatural (archaeology)Social mediaNatural language processingArtificial intelligenceWorld Wide WebHistoryArchaeology

Abstract

fetched live from OpenAlex

Background: Nutritional yeast, an inactive form of Saccharomyces cerevisiae, has recently become increasingly popular as a food supplement and healthy ingredient, especially among individuals following plant-based diets. It is valued for its health benefits and high content of B vitamins, minerals, and protein. Social media has enabled people to share information and personal experiences at an unprecedented level, further amplifying conversations around health and nutrition. With the rise of social media, data mining techniques like natural language processing and machine learning are increasingly used for analyzing the large amounts of information generated on these platforms. Objective: This study aimed to analyze social media data from the United States and France to identify the most frequently discussed topics among nutritional yeast consumers. The objective was to fill gaps in our understanding of the perceptions, experiences, and usage trends related to nutritional yeast. Methods: This study was retrospective, using social media data geolocated in the United States and France, posted by users discussing nutritional yeast between December 2017 and September 2023. Data cleaning and filtering were done using natural language processing methods and specific algorithms. Biterm topic modeling was applied to identify the most frequently discussed topics. Results: A total of 36,642 posts written by 28,069 users discussing nutritional yeast were identified across 1039 publicly available online sources. This included 34,292 posts from the United States (26,154 users across 994 sources) and 2350 from France (n=1915 users across 45 sources). Twitter was the most commonly used platform in both countries, accounting for 39.6% of posts in the United States (13,587/34,292) and 84.3% in France (1982/2350). In the United States, conversations centered around the role of nutritional yeast role as a vegan nutrient source (n=12,345, 36.0%). Several users highlighted its culinary versatility as a natural seasoning (n=8093, 23.6%) and its health and skin benefits (n=6173, 18.0%). In France, discussions frequently focused on nutritional yeast's use in dietary supplement routines in various forms (n=1177, 50.1%), emphasizing its benefits alongside other supplements such as castor oil, particularly noted for effects on nails and hair (n=928, 39.5%). Conclusions: This social media listening study identified the perceptions and preferences of nutritional yeast users in France and the United States. Researchers and health care professionals can reflect on these findings to investigate the potential health benefits of nutritional yeast for specific groups and its long-term impact on different diets and lifestyles. Marketers may also use this information to create customized strategies that better align with the preferences and needs of each market.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.031
Threshold uncertainty score0.061

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.008
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0050.003
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0000.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.040
GPT teacher head0.422
Teacher spread0.382 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR InfodemiologySame topicSocial Media in Health EducationFrench-language works237,207