Prepandemic Antivaccination Websites' COVID-19 Vaccine Behavior: Content Analysis of Archived Websites
Bibliographic record
Abstract
BACKGROUND: The onset of the COVID-19 pandemic and the concurrent development of vaccines offered a rare and somewhat unprecedented opportunity to study antivaccination behavior as it formed over time via the use of archived versions of websites. OBJECTIVE: This study aims to assess how existing antivaccination websites modified their content to address COVID-19 vaccines and pandemic restrictions. METHODS: Using a preexisting collection of 25 antivaccination websites curated by the IvyPlus Web Collection Program prior to the pandemic and crawled every 6 months via Archive-It, we conducted a content analysis to see how these websites acknowledged or ignored COVID-19 vaccines and pandemic restrictions. Websites were assessed for financial behaviors such as having storefronts, mention of COVID-19 vaccines in general or by manufacturer name, references to personal freedom such as masking, safety concerns like side effects, and skepticism of science. RESULTS: The majority of websites addressed COVID-19 vaccines in a negative fashion, with more websites making appeals to personal freedom or expressing skepticism of science than questioning safety. This can potentially be attributed to the lack of available safety data about the vaccines at the time of data collection. Many of the antivaccination websites we evaluated actively sought donations and had a membership option, evidencing these websites have financial motivations and actively build a community around these issues. The content analysis also offered the opportunity to test the viability of archived websites for use in scholarly research. The archived versions of the websites had significant shortcomings, particularly in search functionality, and required supplementation with the live websites. For web archiving to be a viable source of stand-alone content for research, the technology needs to make significant improvements in its capture abilities. CONCLUSIONS: In summary, we found antivaccination websites existing prior to the COVID-19 pandemic largely adapted their messaging to address COVID-19 vaccines with very few sites ignoring the pandemic altogether. This study also demonstrated the timely and significant need for more robust web archiving capabilities as web-based environments become more ephemeral and unstable.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.035 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.006 | 0.005 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".