Exploring Conceptualizations of COVID-19 Risk in Ideologically Distinct Online Communities: A Computational Grounded Theory Analysis
Bibliographic record
Abstract
BACKGROUND: The COVID-19 pandemic has had a profound impact on societies and economies around the globe, and experts warn about the potential for similar crises in the future. Risk communication theories underscore that while the potential for harm is objective, risk perception is a subjective, socially derived interpretation. While there is broad literature on the social construction of risk, fewer studies examine the role of communities-online or offline-in developing and reinforcing distinct interpretations of the same risk event. During COVID-19, online communities emerged as individuals sought to make sense of the ongoing crisis. These communities offer an opportunity to gain important insights into how concerned public collectively interprets risk and create group identities, informing public health strategies. OBJECTIVE: This study aims to, first, explore how online communities with distinct ideologies create and reinforce divergent conceptualizations of risk and, second, identify the role of group identity in shaping the development and communication of risk interpretations in these communities. METHODS: We used computational grounded theory, a multistep approach that includes pattern detection, hypothesis testing, and pattern confirmation to explore interpretations of risk and group identity in about 500,000 comments from the subreddits r/LockdownSkepticism and r/Masks4All. In the pattern detection step of this study, we grouped comments by the post they were made on and then used latent Dirichlet allocation topic modeling to identify 10 topics based on the frequency of term co-occurrence. In the hypothesis refinement step, we conducted a qualitative thematic analysis of 30 posts under each topic using Braun and Clarke's approach. Finally, in the pattern confirmation step, we trained a Word2Vec word embedding model to validate emerging themes from the second step. RESULTS: This study found that Masks4All and LockdownSkepticism both centered risk in their conversations, but with divergent concerns related to the threat of COVID-19. While Masks4All emphasized the threat to health, LockdownSkepticism questioned the necessity of preventive measures and focused on other risks: the threat to the economy, educational disruptions, and social isolation. Group identity was also found to shape collective meanings around risk, as community members in both subreddits affirmed group positions and condemned the outgroup. CONCLUSIONS: This study demonstrated that while both communities were concerned about COVID-19, their perceptions of risk focused on different aspects of the same risk event. This underscores the need for targeted interventions that engage with divergent ideologies and value systems across groups of people.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.017 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".