Monitoring and Identifying Emerging e-Cigarette Brands and Flavors on Twitter: Observational Study
Bibliographic record
Abstract
BACKGROUND: Flavored electronic cigarettes (e-cigarettes) have become very popular in recent years. e-Cigarette users like to share their e-cigarette products and e-cigarette use (vaping) experiences on social media. e-Cigarette marketing and promotions are also prevalent online. OBJECTIVE: This study aims to develop a method to identify new e-cigarette brands and flavors mentioned on Twitter and to monitor e-cigarette brands and flavors mentioned on Twitter from May 2021 to December 2021. METHODS: We collected 1.9 million tweets related to e-cigarettes between May 3, 2021, and December 31, 2021, by using the Twitter streaming application programming interface. Commercial and noncommercial tweets were characterized based on promotion-related keywords. We developed a depletion method to identify new e-cigarette brands by removing the keywords that already existed in the reference data set (Twitter data related to e-cigarettes from May 3, 2021, to August 31, 2021) or our previously identified brand list from the keywords in the target data set (e-cigarette-related Twitter data from September 1, 2021, to December 31, 2021), followed by a manual Google search to identify new e-cigarette brands. To identify new e-cigarette flavors, we constructed a flavor keyword list based on our previously collected e-cigarette flavor names, which were used to identify potential tweet segments that contain at least one of the e-cigarette flavor keywords. Tweets or tweet segments with flavor keywords but not any known flavor names were marked as potential new flavor candidates, which were further verified by a web-based search. The longitudinal trends in the number of tweets mentioning e-cigarette brands and flavors were examined in both commercial and noncommercial tweets. RESULTS: Through our developed methods, we identified 34 new e-cigarette brands and 97 new e-cigarette flavors from commercial tweets as well as 56 new e-cigarette brands and 164 new e-cigarette flavors from noncommercial tweets. The longitudinal trend of the e-cigarette brands showed that JUUL was the most popular e-cigarette brand mentioned on Twitter; however, there was a decreasing trend in the mention of JUUL over time on Twitter. Menthol flavor was the most popular e-cigarette flavor mentioned in the commercial tweets, whereas mango flavor was the most popular e-cigarette flavor mentioned in the noncommercial tweets during our study period. CONCLUSIONS: Our proposed methods can successfully identify new e-cigarette brands and flavors mentioned on Twitter. Twitter data can be used for monitoring the dynamic changes in the popularity of e-cigarette brands and flavors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".