The Role of Social Media in Health Misinformation and Disinformation During the COVID-19 Pandemic: Bibliometric Analysis
Bibliographic record
Abstract
BACKGROUND: The use of social media platforms to seek information continues to increase. Social media platforms can be used to disseminate important information to people worldwide instantaneously. However, their viral nature also makes it easy to share misinformation, disinformation, unverified information, and fake news. The unprecedented reliance on social media platforms to seek information during the COVID-19 pandemic was accompanied by increased incidents of misinformation and disinformation. Consequently, there was an increase in the number of scientific publications related to the role of social media in disseminating health misinformation and disinformation at the height of the COVID-19 pandemic. Health misinformation and disinformation, especially in periods of global public health disasters, can lead to the erosion of trust in policy makers at best and fatal consequences at worst. OBJECTIVE: This paper reports a bibliometric analysis aimed at investigating the evolution of research publications related to the role of social media as a driver of health misinformation and disinformation since the start of the COVID-19 pandemic. Additionally, this study aimed to identify the top trending keywords, niche topics, authors, and publishers for publishing papers related to the current research, as well as the global collaboration between authors on topics related to the role of social media in health misinformation and disinformation since the start of the COVID-19 pandemic. METHODS: The Scopus database was accessed on June 8, 2023, using a combination of Medical Subject Heading and author-defined terms to create the following search phrases that targeted the title, abstract, and keyword fields: ("Health*" OR "Medical") AND ("Misinformation" OR "Disinformation" OR "Fake News") AND ("Social media" OR "Twitter" OR "Facebook" OR "YouTube" OR "WhatsApp" OR "Instagram" OR "TikTok") AND ("Pandemic*" OR "Corona*" OR "Covid*"). A total of 943 research papers published between 2020 and June 2023 were analyzed using Microsoft Excel (Microsoft Corporation), VOSviewer (Centre for Science and Technology Studies, Leiden University), and the Biblioshiny package in Bibliometrix (K-Synth Srl) for RStudio (Posit, PBC). RESULTS: The highest number of publications was from 2022 (387/943, 41%). Most publications (725/943, 76.9%) were articles. JMIR published the most research papers (54/943, 5.7%). Authors from the United States collaborated the most, with 311 coauthored research papers. The keywords "Covid-19," "social media," and "misinformation" were the top 3 trending keywords, whereas "learning systems," "learning models," and "learning algorithms" were revealed as the niche topics on the role of social media in health misinformation and disinformation during the COVID-19 outbreak. CONCLUSIONS: Collaborations between authors can increase their productivity and citation counts. Niche topics such as "learning systems," "learning models," and "learning algorithms" could be exploited by researchers in future studies to analyze the influence of social media on health misinformation and disinformation during periods of global public health emergencies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.127 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.216 | 0.332 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.008 | 0.006 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".