Utility of Facebook’s Social Connectedness Index in Modeling COVID-19 Spread: Exponential Random Graph Modeling Study
Bibliographic record
Abstract
BACKGROUND: The COVID-19 (the disease caused by the SARS-CoV-2 virus) pandemic has underscored the need for additional data, tools, and methods that can be used to combat emerging and existing public health concerns. Since March 2020, there has been substantial interest in using social media data to both understand and intervene in the pandemic. Researchers from many disciplines have recently found a relationship between COVID-19 and a new data set from Facebook called the Social Connectedness Index (SCI). OBJECTIVE: Building off this work, we seek to use the SCI to examine how social similarity of Missouri counties could explain similarities of COVID-19 cases over time. Additionally, we aim to add to the body of literature on the utility of the SCI by using a novel modeling technique. METHODS: In September 2020, we conducted this cross-sectional study using publicly available data to test the association between the SCI and COVID-19 spread in Missouri using exponential random graph models, which model relational data, and the outcome variable must be binary, representing the presence or absence of a relationship. In our model, this was the presence or absence of a highly correlated COVID-19 case count trajectory between two given counties in Missouri. Covariates included each county's total population, percent rurality, and distance between each county pair. RESULTS: We found that all covariates were significantly associated with two counties having highly correlated COVID-19 case count trajectories. As the log of a county's total population increased, the odds of two counties having highly correlated COVID-19 case count trajectories increased by 66% (odds ratio [OR] 1.66, 95% CI 1.43-1.92). As the percent of a county classified as rural increased, the odds of two counties having highly correlated COVID-19 case count trajectories increased by 1% (OR 1.01, 95% CI 1.00-1.01). As the distance (in miles) between two counties increased, the odds of two counties having highly correlated COVID-19 case count trajectories decreased by 43% (OR 0.57, 95% CI 0.43-0.77). Lastly, as the log of the SCI between two Missouri counties increased, the odds of those two counties having highly correlated COVID-19 case count trajectories significantly increased by 17% (OR 1.17, 95% CI 1.09-1.26). CONCLUSIONS: These results could suggest that two counties with a greater likelihood of sharing Facebook friendships means residents of those counties have a higher likelihood of sharing similar belief systems, in particular as they relate to COVID-19 and public health practices. Another possibility is that the SCI is picking up travel or movement data among county residents. This suggests the SCI is capturing a unique phenomenon relevant to COVID-19 and that it may be worth adding to other COVID-19 models. Additional research is needed to better understand what the SCI is capturing practically and what it means for public health policies and prevention practices.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".