Twitter Conversations About Pancreatic Cancer by Health Care Providers and the General Public: Thematic Analysis
Bibliographic record
Abstract
BACKGROUND: There is a growing interest in the pattern of consumption of health-related information on social media platforms. OBJECTIVE: We evaluated the content of discussions around pancreatic cancer on Twitter to identify subtopics of greatest interest to health care providers and the general public. METHODS: We used an online analytical tool (Creation Pinpoint) to quantify Twitter mentions (tweets and retweets) related to pancreatic cancer between January 2018 and December 2019. Keywords, hashtags, word combinations, and phrases were used to identify mentions. Health care provider profiles were identified using machine learning and then verified by a human analyst. Remaining user profiles were classified as belonging to the general public. Data from conversations were stratified qualitatively into 5 domains: (1) prevention, (2) survivorship, (3) treatment, (4) research, and (5) policy. We compared the themes of conversations initiated by health care providers and the general public and analyzed the impact of the Pancreatic Cancer Awareness Month and announcements by public figures of pancreatic cancer diagnoses on the overall volume of conversations. RESULTS: Out of 1,258,028 mentions of pancreatic cancer, 313,668 unique mentions were classified into the 5 domains. We found that health care providers most commonly discussed pancreatic cancer research (10,640/27,031 mentions, 39.4%), while the general public most commonly discussed treatment (154,484/307,449 mentions, 50.2%). Health care providers were found to be more likely to initiate conversations related to research (odds ratio [OR] 1.75, 95% CI 1.70-1.79, P<.001) and prevention (OR 1.49, 95% CI 1.41-1.57, P<.001) whereas the general public took the lead in the domains of treatment (OR 1.63, 95% CI 1.58-1.69, P<.001) and survivorship (OR 1.17, 95% CI 1.13-1.21, P<.001). Pancreatic Cancer Awareness Month did not increase the number of mentions by health care providers in any of the 5 domains, but general public mentions increased temporarily in all domains except prevention and policy. Health care provider mentions did not increase with announcements by public figures of pancreatic cancer diagnoses. After Alex Trebek, host of the television show Jeopardy, received his diagnosis, general public mentions of survivorship increased, while Justice Ruth Bader Ginsburg's diagnosis increased conversations on treatment. CONCLUSIONS: Health care provider conversations on Twitter are not aligned with the general public. Pancreatic Cancer Awareness Month temporarily increased general public conversations about treatment, research, and survivorship, but not prevention or policy. Future studies are needed to understand how conversations on social media platforms can be leveraged to increase health care awareness among the general public.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.002 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".