Mapping Vaccine Sentiment by Analyzing Spanish-Language Social Media Posts and Survey-Based Public Opinion: Dual Methods Study
Bibliographic record
Abstract
Background: The internet and social media have been considered useful platforms for obtaining health information. However, critical and erroneous content about vaccines on social media has been associated with vaccination delays and refusal. Objective: This study aimed to examine how social networks influence access to and perceptions of vaccine-related information. We sought to (1) quantify the proportion of individuals engaging with vaccine-related content on social media and to characterize their demographic and behavioral profiles through an internet-based population survey conducted in Spain and (2) to analyze vaccine-related sentiments and opinions in Spanish and Catalan posts on X (X Corp [formerly Twitter, Inc] and geolocate them using artificial intelligence. Methods: Two complementary methodologies were applied. First, an observational study was conducted via a self-administered internet-based questionnaire among adults in Spain in 2021. Second, we analyzed Spanish- and Catalan-language posts from X, collected between March and December 2021. Sentiment analysis was performed using a workflow developed in Orange Data Mining (Bioinformatics Laboratory, Faculty of Computer and Information Science, University of Ljubljana). Geolocation was based on user-defined locations and visualized using Microsoft Power Business Intelligence. Social network analysis was conducted with NodeXL Pro (Social Media Research Foundation) to identify and characterize the 5 largest user communities discussing vaccines. Although based on independent data sources, the 2 approaches provided complementary methodological insights. Results: Among the 1312 respondents in the survey, 85.7% (1124/1312) stated that they were regular social network users, and 66% (850/1287) reported having encountered antivaccine information on social networks. Of these, 24.3% (205/845) experienced doubts about receiving recommended vaccines, and out of those with doubts, 13.3% (27/203) refused at least 1 vaccine proposed by a health care professional. A total of 479,734 Spanish and Catalan posts on X were analyzed, with 54.44% (n=261,183) posts classified as negative, 28.18% (n=135,194) as neutral, and 17.37% (n=83,357) as positive. Sentiment varied across regions, with more negative posts appearing to derive from South America, with a mix in Europe and more positive posts in North America. Analysis of the topic words and key themes allowed the grouping of the predominant themes of the 5 study groups, which were (1) vaccination efforts during the COVID-19 pandemic, (2) issues of vaccine theft and struggles in managing and securing the vaccine supply, (3) campaigns in the State of Mexico, (4) vaccination efforts for older adults, and (5) the vaccination campaign in Colombia to combat COVID-19. Conclusions: High proportions of exposure to antivaccine content were reported by the surveyed population. Sentiment analysis and geolocation of posts on the social network X suggested a notable presence of Spanish-language posts categorized as negative, predominantly from South America. The thematic analysis of conversations on X may provide valuable insights into the population's opinions about vaccines.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".