Misinformation and Profitability of Hepatitis B Virus Claims on Instagram: Formative Cross-Sectional Study
Bibliographic record
Abstract
Background The internet is increasingly used to find health information, which often contains misinformation. Instagram is a likely source of health information online for many adults worldwide, given that there are more than 2 billion worldwide users. To date, no studies have documented the characteristics of hepatitis B virus (HBV) claims, information accuracy, engagement with, and profitability of HBV information on Instagram. Objective We aimed to document the characteristics, accuracy, engagement, and profitability of HBV misinformation on Instagram. Methods In this cross-sectional formative study, 2 research members searched for publicly available Instagram posts using the terms “hepatitis b” and “hep b” and manually extracted data from the most popular posts and user profiles for each term from December 2021 to January 2022 at varying times of the day and days of the week. We applied an existing and validated health misinformation codebook, adapted for this topic, to 103 posts for 58 variables, including post characteristics, types of HBV claims (eg, treatment, prevention, and cure), accuracy of information (misinformation vs accurate, coded by hepatology clinicians), engagement (number of likes), and profitability (yes or no). We calculated descriptive statistics and applied chi-square, Fisher exact, and z tests to compare posts with certain characteristics, claims, and engagement by accuracy and profitability in Stata (version 18.0) with significance set at an α of .05. Results Of the full sample, most posts had accurate (79/103, 76.7%) versus inaccurate (24/103, 23.3%) information about HBV. Among posts with claims about HBV treatment (18/103, 17.5%), there were more posts that had misinformation than accurate posts (55.6% vs 44.4%; χ²1=12.7; P<.001). Similarly, there were higher proportions of posts with misinformation compared to posts with accurate information about cures (n=12, 75% vs 25%; Fisher P<.001), natural remedies (n=13, 92.3% vs 7.7%; Fisher P<.001), symptoms (n=15, 60% vs 40%; χ²1=13.2; P<.001), and censorship conspiracies (n=9, 66.7% vs 33.3%; Fisher P=.005) related to HBV. Compared to posts with accurate information, posts with misinformation had more likes on average (mean 1459.2, SD 1458.8-1459.6 vs mean 941.8, SD 941.6-942.0; z=−517.4; P<.001). Significantly more posts with misinformation were for profit (39.5% vs 13.8%; χ²1=8.8; P=.003) than accurate posts. Conclusions HBV misinformation had more engagement than accurate information on Instagram and was more likely to be for-profit than accurate information. HBV misinformation may spread more easily than accurate information, meaning people searching for HBV on Instagram may encounter false, profit-driven claims that could affect health behaviors. Our focus on visual social media misinformation is innovative, as is our use of Instagram, an understudied platform. More research is needed to estimate the prevalence of HBV misinformation and its influence on health beliefs, behaviors, and outcomes. Improving media literacy may help reduce the influence of HBV misinformation online.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".