Misinformation in Italian Online Mental Health Communities During the COVID-19 Pandemic: Protocol for a Content Analysis Study
Bibliographic record
Abstract
BACKGROUND: Social media platforms are widely used by people suffering from mental illnesses to cope with their conditions. One modality of coping with these conditions is navigating online communities where people can receive emotional support and informational advice. Benefits have been documented in terms of impact on health outcomes. However, the pitfalls are still unknown, as not all content is necessarily helpful or correct. Furthermore, the advent of the COVID-19 pandemic and related problems, such as worsening mental health symptoms, the dissemination of conspiracy narratives, and medical distrust, may have impacted these online communities. The situation in Italy is of particular interest, being the first Western country to experience a nationwide lockdown. Particularly during this challenging time, the beneficial role of community moderators with professional mental health expertise needs to be investigated in terms of uncovering misleading information and regulating communities. OBJECTIVE: The aim of the proposed study is to investigate the potentially harmful content found in online communities for mental health symptoms in the Italian language. Besides descriptive information about the content that posts and comments address, this study aims to analyze the content from two viewpoints. The first one compares expert-led and peer-led communities, focusing on differences in misinformation. The second one unravels the impact of the COVID-19 pandemic, not by merely investigating differences in topics but also by investigating the needs expressed by community members. METHODS: A codebook for the content analysis of Facebook communities has been developed, and a content analysis will be conducted on bundles of posts. Among 14 Facebook groups that were interested in participating in this study, two groups were selected for analysis: one was being moderated by a health professional (n=12,058 members) and one was led by peers (n=5598 members). Utterances from 3 consecutive calendar years will be studied by comparing the months from before the pandemic, the months during the height of the pandemic, and the months during the postpandemic phase (2019-2021). This method permits the identification of different types of misinformation and the context in which they emerge. Ethical approval was obtained by the Università della Svizzera italiana ethics committee. RESULTS: The usability of the codebook was demonstrated with a pretest. Subsequently, 144 threads (1534 utterances) were coded by the two coders. Intercoder reliability was calculated on 293 units (19.10% of the total sample; Krippendorff α=.94, range .72-1). Aside from a few analyses comparing bundles, individual utterances will constitute the unit of analysis in most cases. CONCLUSIONS: This content analysis will identify deleterious content found in online mental health support groups, the potential role of moderators in uncovering misleading information, and the impact of COVID-19 on the content. INTERNATIONAL REGISTERED REPORT IDENTIFIER (IRRID): PRR1-10.2196/35347.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.004 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".