MétaCan
Menu
Back to cohort
Record W4414347749 · doi:10.2196/76737

Response to the Netflix Docuseries “Big Vape: The Rise and Fall of JUUL”: Mixed Methods Analysis of YouTube Comments Using Qualitative Coding and Topic Modeling

2025· article· en· W4414347749 on OpenAlexvenueno aff
Beth L. Hoffman, Arpita Tripathi, Ariel Shensa, Pengyue Dou, Piper Narendorf, Nishi Hundi, Jaime E. Sidani

Bibliographic record

VenueJMIR Formative Research · 2025
Typearticle
Languageen
FieldSocial Sciences
TopicMisinformation and Its Impacts
Canadian institutionsnot available
FundersNational Institute on Minority Health and Health Disparities
KeywordsMisinformationSocial mediaQualitative analysisTopic modelCoding (social sciences)PerceptionHealth communication

Abstract

fetched live from OpenAlex

Background: On October 11, 2023, Netflix released the docuseries "Big Vape: The Rise and Fall of JUUL," which chronicled the founding of JUUL, its rise in popularity among youth, and the subsequent public backlash. The official Netflix YouTube channel posted a trailer promoting the docuseries and an official clip from the docuseries. Recent studies have demonstrated the utility of using comments posted under YouTube videos to analyze reactions to the content and discourse around the health topics explored in the video. Objective: This study aimed to (1) systematically characterize nicotine and tobacco product (NTP)-related comments and replies posted in response to the docuseries trailer and video clip and (2) explore integration of automated topic modeling techniques with traditional human-generated qualitative coding. Methods: We extracted all comments and replies on the aforementioned YouTube clips 1 month after the docuseries' release (N=532). Research assistants manually double-coded the comments using a systematically developed codebook that assessed for NTP sentiment (pro-NTP, anti-NTP, complex sentiment, or no sentiment) and the presence or absence of specific electronic cigarette (e-cigarette)-related content. Given the substantial amount of comments coded as potential misinformation during the coding process, we conducted an in-depth qualitative content analysis of all comments coded as potential misinformation. Simultaneously, we used word clustering techniques including structural topic modeling to identify the overarching topics. Results: Of the 73.8% ( 393/532) relevant comments, 63.6% (250/393) expressed NTP sentiment with 42.8% of these (107/250) expressing pro-NTP sentiment and 18.4% (46/250) expressing complex sentiment. The most frequent content category was potential misinformation (27.5%, 108/393). These 108 comments contained 152 individual pieces of misinformation that were broadly grouped within 6 themes with various numbers of subthemes; the most frequent misinformation theme was that e-cigarette use is completely safe or much safer than smoking (n=80). Other frequently occurring content categories included e-cigarette use is safer than smoking (17.6%, 69/393), and personal experience using e-cigarettes or JUUL (15.5%, 61/393). For topic modeling, we identified 9 topics that we qualitatively assigned into 4 thematic categories: comparisons with other drugs, mentions of government and pharma companies, role of media and parents, and harms associated with nicotine and tobacco products. Conclusions: To the best of our knowledge, this is the first study to examine viewer reactions to the docuseries about JUUL. Our analysis of YouTube comments offers insight into current sentiment and misinformation regarding NTPs and highlights the potential utility of using mixed methods to analyze NTP-related social media data, and the benefits of integrating computational and human qualitative research to analyze social media perceptions of e-cigarettes. Public health professionals can use our findings to help develop tailored health communication messages to address common sentiment and misconceptions related to JUUL, other e-cigarette products, and new NTP products.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.022
metaresearch head score (Gemma)0.072
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: Qualitative
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.022
Threshold uncertainty score0.116

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0220.072
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0040.003
Science and technology studies0.0030.002
Scholarly communication0.0020.002
Open science0.0010.003
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.269
GPT teacher head0.598
Teacher spread0.329 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designQualitative
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJMIR Formative ResearchSame topicMisinformation and Its ImpactsFrench-language works237,207