Analyzing Trends of Loneliness Through Large-Scale Analysis of Social Media Postings: Observational Study
Bibliographic record
Abstract
BACKGROUND: Loneliness has become a public health problem described as an epidemic, and it has been argued that digital behavior such as social media posting affects loneliness. OBJECTIVE: The aim of this study is to expand knowledge of the determinants of loneliness by investigating online postings in a social media forum devoted to loneliness. Specifically, this study aims to analyze the temporal trends in loneliness and their associations with topics of interest, especially with those related to mental health determinants. METHODS: We collected a total of 19,668 postings from 11,054 users in the loneliness forum on Reddit. We asked seven crowdsourced workers to imagine themselves as writing 1 of 236 randomly chosen posts and to answer the short-form UCLA Loneliness Scale. After showing that these postings could provide an assessment of loneliness, we built a predictive model for loneliness scores based on the posts' text and applied it to all collected postings. We then analyzed trends in loneliness postings over time and their correlations with other topics of interest related to mental health determinants. RESULTS: We found that crowdsourced workers can estimate loneliness (interclass correlation=0.19) and that predictive models are correlated with reported loneliness scores (Pearson r=0.38). Our results show that increases in loneliness are strongly associated with postings to a suicidality-related forum (hazard ratio 1.19) and to forums associated with other detrimental behaviors such as depression and illicit drug use. Clustering demonstrates that people who are lonely come from diverse demographics and from a variety of interests. CONCLUSIONS: The results demonstrate that it is possible for unrelated individuals to assess people's social media postings for loneliness. Moreover, our findings show the multidimensional nature of online loneliness and its correlated behaviors. Our study shows the advantages of studying a hard-to-reach population through social media and suggests new directions for future studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".