Mining Physicians’ Opinions on Social Media to Obtain Insights Into COVID-19: Mixed Methods Analysis
Bibliographic record
Abstract
BACKGROUND: The coronavirus disease (COVID-19) pandemic is considered to be the most daunting public health challenge in decades. With no effective treatments and with time needed to develop a vaccine, alternative approaches are being used to control this pandemic. OBJECTIVE: The objective of this paper was to identify topics, opinions, and recommendations about the COVID-19 pandemic discussed by medical professionals on the Twitter social medial platform. METHODS: Using a mixed methods approach blending the capabilities of social media analytics and qualitative analysis, we analyzed COVID-19-related tweets posted by medical professionals and examined their content. We used qualitative analysis to explore the collected data to identify relevant tweets and uncover important concepts about the pandemic using qualitative coding. Unsupervised and supervised machine learning techniques and text analysis were used to identify topics and opinions. RESULTS: Data were collected from 119 medical professionals on Twitter about the coronavirus pandemic. A total of 10,096 English tweets were collected from the identified medical professionals between December 1, 2019 and April 1, 2020. We identified eight topics, namely actions and recommendations, fighting misinformation, information and knowledge, the health care system, symptoms and illness, immunity, testing, and infection and transmission. The tweets mainly focused on needed actions and recommendations (2827/10,096, 28%) to control the pandemic. Many tweets warned about misleading information (2019/10,096, 20%) that could lead to infection of more people with the virus. Other tweets discussed general knowledge and information (911/10,096, 9%) about the virus as well as concerns about the health care systems and workers (909/10,096, 9%). The remaining tweets discussed information about symptoms associated with COVID-19 (810/10,096, 8%), immunity (707/10,096, 7%), testing (605/10,096, 6%), and virus infection and transmission (503/10,096, 5%). CONCLUSIONS: Our findings indicate that Twitter and social media platforms can help identify important and useful knowledge shared by medical professionals during a pandemic.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.050 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".