Linking Individual-Level Facebook Posts With Psychological and Health Data in an Epidemiological Cohort: Feasibility Study
Bibliographic record
Abstract
BACKGROUND: Psychological factors (eg, depression) and related biological and behavioral responses are associated with numerous physical health outcomes. Most research in this area relies on self-reported assessments of psychological factors, which are difficult to scale because they may be expensive and time-consuming. Investigators are increasingly interested in using social media as a novel and convenient platform for obtaining information rapidly in large populations. OBJECTIVE: We evaluated the feasibility of obtaining Facebook data from a large ongoing cohort study of midlife and older women, which may be used to assess psychological functioning efficiently with low cost. METHODS: This study was conducted with participants in the Nurses' Health Study II (NHSII), which was initiated in 1989 with biennial follow-ups. Facebook does not share data readily; therefore, we developed procedures to enable women to download and transfer their Facebook data to cohort servers (for linkage with other study data they have provided). Since privacy is a critical concern when collecting individual-level data, we partnered with a third-party software developer, Digi.me, to enable participants to obtain their own Facebook data and to send it securely to our research team. In 2020, we invited a subset of the 18,519 NHSII participants (aged 56-73 years) via email to participate. Women were selected if they reported on the 2017-2018 questionnaire that they regularly posted on Facebook and were still active cohort participants. We included an exit survey for those who chose not to participate in order to gauge the reasons for nonparticipation. RESULTS: We invited 309 women to participate. Few women signed the consent form (n=52), and only 3 used the Digi.me app to download and transfer their Facebook data. This low participation rate was observed despite modifying our protocol between waves of recruitment, including by (1) excluding active health care workers, who might be less available to participate due to the pandemic, (2) developing a Frequently Asked Questions factsheet to provide more information regarding the protocol, and (3) simplifying the instructions for using the Digi.me app. On our exit survey, the reasons most commonly reported for not participating were concerns regarding data privacy and hesitation sharing personal Facebook posts. The low participation rate suggests that obtaining individual-level Facebook data in a cohort of middle-aged and older women may be challenging. CONCLUSIONS: In this cohort of midlife and older women who were actively participating for over three decades, we were largely unable to obtain permission to access individual-level data from participants' Facebook accounts. Despite working with a third-party developer to customize an app to implement safeguards for privacy, data privacy remained a key concern in these women. Future studies aiming to leverage individual-level social media data should explore alternate populations or means of sharing social media data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.052 | 0.054 |
| Meta-epidemiology (narrow) | 0.001 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.004 | 0.002 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".