Beyond Numbers: Opportunities and Challenges Using Qualitative Methods in Social Media Studies of PrEP Discourse
Bibliographic record
Abstract
Using pre-exposure prophylaxis (PrEP) discourse as a case study, we examine qualitative and multi-method approaches to analyzing health communication on social media platforms. These platforms have become crucial spaces for health promotion and community discourse, particularly on HIV prevention. Our review of 23 PrEP-focused studies that used social media data documents the different strengths of both methodologies. Qualitative analyses were strong at capturing how platform architecture shaped discourse quality: Reddit's forum structure enabled deeper narratives while X's (formerly Twitter) character constrained discussions. Algorithm-driven platforms' short-form video formats like TikTok emerged as unique media for health communication, enabling creative expression but limiting discourse depth. Multi-method studies dominated the literature (n = 15) yet struggled to meaningfully integrate quantitative metrics with qualitative insights. Frameworks are required that better combine computational approaches with qualitative analysis, while accounting for key challenges of social media data: the transient nature of data, context preservation, and post-authenticity validation.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".