Using Patient-Generated Health Data From Twitter to Identify, Engage, and Recruit Cancer Survivors in Clinical Trials in Los Angeles County: Evaluation of a Feasibility Study
Bibliographic record
Abstract
BACKGROUND: Failure to find and attract clinical trial participants remains a persistent barrier to clinical research. Researchers increasingly complement recruitment methods with social media-based methods. We hypothesized that user-generated data from cancer survivors and their family members and friends on the social network Twitter could be used to identify, engage, and recruit cancer survivors for cancer trials. OBJECTIVE: This pilot study aims to examine the feasibility of using user-reported health data from cancer survivors and family members and friends on Twitter in Los Angeles (LA) County to enhance clinical trial recruitment. We focus on 6 cancer conditions (breast cancer, colon cancer, kidney cancer, lymphoma, lung cancer, and prostate cancer). METHODS: The social media intervention involved monitoring cancer-specific posts about the 6 cancer conditions by Twitter users in LA County to identify cancer survivors and their family members and friends and contacting eligible Twitter users with information about open cancer trials at the University of Southern California (USC) Norris Comprehensive Cancer Center. We reviewed both retrospective and prospective data published by Twitter users in LA County between July 28, 2017, and November 29, 2018. The study enrolled 124 open clinical trials at USC Norris. We used descriptive statistics to report the proportion of Twitter users who were identified, engaged, and enrolled. RESULTS: We analyzed 107,424 Twitter posts in English by 25,032 unique Twitter users in LA County for the 6 cancer conditions. We identified and contacted 1.73% (434/25,032) of eligible Twitter users (127/434, 29.3% cancer survivors; 305/434, 70.3% family members and friends; and 2/434, 0.5% Twitter users were excluded). Of them, 51.4% (223/434) were female and approximately one-third were male. About one-fifth were people of color, whereas most of them were White. Approximately one-fifth (85/434, 19.6%) engaged with the outreach messages (cancer survivors: 33/85, 38% and family members and friends: 52/85, 61%). Of those who engaged with the messages, one-fourth were male, the majority were female, and approximately one-fifth were people of color, whereas the majority were White. Approximately 12% (10/85) of the contacted users requested more information and 40% (4/10) set up a prescreening. Two eligible candidates were transferred to USC Norris for further screening, but neither was enrolled. CONCLUSIONS: Our findings demonstrate the potential of identifying and engaging cancer survivors and their family members and friends on Twitter. Optimization of downstream recruitment efforts such as screening for digital populations on social media may be required. Future research could test the feasibility of the approach for other diseases, locations, languages, social media platforms, and types of research involvement (eg, survey research). Computer science methods could help to scale up the analysis of larger data sets to support more rigorous testing of the intervention. TRIAL REGISTRATION: ClinicalTrials.gov NCT03408561; https://clinicaltrials.gov/ct2/show/NCT03408561.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.062 | 0.095 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.003 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".