Association Between What People Learned About COVID-19 Using Web Searches and Their Behavior Toward Public Health Guidelines: Empirical Infodemiology Study
Bibliographic record
Abstract
BACKGROUND: The use of the internet and web-based platforms to obtain public health information and manage health-related issues has become widespread in this digital age. The practice is so pervasive that the first reaction to obtaining health information is to "Google it." As SARS-CoV-2 broke out in Wuhan, China, in December 2019 and quickly spread worldwide, people flocked to the internet to learn about the novel coronavirus and the disease, COVID-19. Lagging responses by governments and public health agencies to prioritize the dissemination of information about the coronavirus outbreak through the internet and the World Wide Web and to build trust gave room for others to quickly populate social media, online blogs, news outlets, and websites with misinformation and conspiracy theories about the COVID-19 pandemic, resulting in people's deviant behaviors toward public health safety measures. OBJECTIVE: The goals of this study were to determine what people learned about the COVID-19 pandemic through web searches, examine any association between what people learned about COVID-19 and behavior toward public health guidelines, and analyze the impact of misinformation and conspiracy theories about the COVID-19 pandemic on people's behavior toward public health measures. METHODS: This infodemiology study used Google Trends' worldwide search index, covering the first 6 months after the SARS-CoV-2 outbreak (January 1 to June 30, 2020) when the public scrambled for information about the pandemic. Data analysis employed statistical trends, correlation and regression, principal component analysis (PCA), and predictive models. RESULTS: The PCA identified two latent variables comprising past coronavirus epidemics (pastCoVepidemics: keywords that address previous epidemics) and the ongoing COVID-19 pandemic (presCoVpandemic: keywords that explain the ongoing pandemic). Both principal components were used significantly to learn about SARS-CoV-2 and COVID-19 and explained 88.78% of the variability. Three principal components fuelled misinformation about COVID-19: misinformation (keywords "biological weapon," "virus hoax," "common cold," "COVID-19 hoax," and "China virus"), conspiracy theory 1 (ConspTheory1; keyword "5G" or "@5G"), and conspiracy theory 2 (ConspTheory2; keyword "ingest bleach"). These principal components explained 84.85% of the variability. The principal components represent two measurements of public health safety guidelines-public health measures 1 (PubHealthMes1; keywords "social distancing," "wash hands," "isolation," and "quarantine") and public health measures 2 (PubHealthMes2; keyword "wear mask")-which explained 84.7% of the variability. Based on the PCA results and the log-linear and predictive models, ConspTheory1 (keyword "@5G") was identified as a predictor of people's behavior toward public health measures (PubHealthMes2). Although correlations of misinformation (keywords "COVID-19," "hoax," "virus hoax," "common cold," and more) and ConspTheory2 (keyword "ingest bleach") with PubHealthMes1 (keywords "social distancing," "hand wash," "isolation," and more) were r=0.83 and r=-0.11, respectively, neither was statistically significant (P=.27 and P=.13, respectively). CONCLUSIONS: Several studies focused on the impacts of social media and related platforms on the spreading of misinformation and conspiracy theories. This study provides the first empirical evidence to the mainly anecdotal discourse on the use of web searches to learn about SARS-CoV-2 and COVID-19.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.004 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.006 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".