Linking Individual-Level Facebook Posts With Psychological and Health Data in an Epidemiological Cohort: Feasibility Study
Notice bibliographique
Résumé
BACKGROUND: Psychological factors (eg, depression) and related biological and behavioral responses are associated with numerous physical health outcomes. Most research in this area relies on self-reported assessments of psychological factors, which are difficult to scale because they may be expensive and time-consuming. Investigators are increasingly interested in using social media as a novel and convenient platform for obtaining information rapidly in large populations. OBJECTIVE: We evaluated the feasibility of obtaining Facebook data from a large ongoing cohort study of midlife and older women, which may be used to assess psychological functioning efficiently with low cost. METHODS: This study was conducted with participants in the Nurses' Health Study II (NHSII), which was initiated in 1989 with biennial follow-ups. Facebook does not share data readily; therefore, we developed procedures to enable women to download and transfer their Facebook data to cohort servers (for linkage with other study data they have provided). Since privacy is a critical concern when collecting individual-level data, we partnered with a third-party software developer, Digi.me, to enable participants to obtain their own Facebook data and to send it securely to our research team. In 2020, we invited a subset of the 18,519 NHSII participants (aged 56-73 years) via email to participate. Women were selected if they reported on the 2017-2018 questionnaire that they regularly posted on Facebook and were still active cohort participants. We included an exit survey for those who chose not to participate in order to gauge the reasons for nonparticipation. RESULTS: We invited 309 women to participate. Few women signed the consent form (n=52), and only 3 used the Digi.me app to download and transfer their Facebook data. This low participation rate was observed despite modifying our protocol between waves of recruitment, including by (1) excluding active health care workers, who might be less available to participate due to the pandemic, (2) developing a Frequently Asked Questions factsheet to provide more information regarding the protocol, and (3) simplifying the instructions for using the Digi.me app. On our exit survey, the reasons most commonly reported for not participating were concerns regarding data privacy and hesitation sharing personal Facebook posts. The low participation rate suggests that obtaining individual-level Facebook data in a cohort of middle-aged and older women may be challenging. CONCLUSIONS: In this cohort of midlife and older women who were actively participating for over three decades, we were largely unable to obtain permission to access individual-level data from participants' Facebook accounts. Despite working with a third-party developer to customize an app to implement safeguards for privacy, data privacy remained a key concern in these women. Future studies aiming to leverage individual-level social media data should explore alternate populations or means of sharing social media data.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,052 | 0,054 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,002 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,004 | 0,002 |
| Communication savante | 0,002 | 0,003 |
| Science ouverte | 0,002 | 0,004 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».