Crowdsourcing Adverse Events Associated With Monoclonal Antibodies Targeting Calcitonin Gene–Related Peptide Signaling for Migraine Prevention: Natural Language Processing Analysis of Social Media
Bibliographic record
Abstract
BACKGROUND: Clinical trials demonstrate the efficacy and tolerability of medications targeting calcitonin gene-related peptide (CGRP) signaling for migraine prevention. However, these trials may not accurately reflect the real-world experiences of more diverse and heterogeneous patient populations, who often have higher disease burden and more comorbidities. Therefore, postmarketing safety surveillance is warranted. Regulatory organizations encourage marketing authorization holders to screen digital media for suspected adverse reactions, applying the same requirements as for spontaneous reports. Real-world data from social media platforms constitute a potential venue to capture diverse patient experiences and help detect treatment-related adverse events. However, while social media holds promise for this purpose, its use in pharmacovigilance is still in its early stages. Computational linguistics, which involves the automatic manipulation and quantitative analysis of oral or written language, offers a potential method for exploring this content. OBJECTIVE: This study aims to characterize adverse events related to monoclonal antibodies targeting CGRP signaling on Reddit, a large online social media forum, by using computational linguistics. METHODS: We examined differences in word frequencies from medication-related posts on the Reddit subforum r/Migraine over a 10-year period (2010-2020) using computational linguistics. The study had 2 phases: a validation phase and an application phase. In the validation phase, we compared posts about propranolol and topiramate, as well as posts about each medication against randomly selected posts, to identify known and expected adverse events. In the application phase, we analyzed posts discussing 2 monoclonal antibodies targeting CGRP signaling-erenumab and fremanezumab-to identify potential adverse events for these medications. RESULTS: From 22,467 Reddit r/Migraine posts, we extracted 402 (2%) propranolol posts, 1423 (6.33%) topiramate posts, 468 (2.08%) erenumab posts, and 73 (0.32%) fremanezumab posts. Comparing topiramate against propranolol identified several expected adverse events, for example, "appetite," "weight," "taste," "foggy," "forgetful," and "dizziness." Comparing erenumab against a random selection of terms identified "constipation" as a recurring keyword. Comparing erenumab against fremanezumab identified "constipation," "depression," "vomiting," and "muscle" as keywords. No adverse events were identified for fremanezumab. CONCLUSIONS: The validation phase of our study accurately identified common adverse events for oral migraine preventive medications. For example, typical adverse events such as "appetite" and "dizziness" were mentioned in posts about topiramate. When we applied this methodology to monoclonal antibodies targeting CGRP or its receptor-fremanezumab and erenumab, respectively-we found no definite adverse events for fremanezumab. However, notable flagged words for erenumab included "constipation," "depression," and "vomiting." In conclusion, computational linguistics applied to social media may help identify potential adverse events for novel therapeutics. While social media data show promise for pharmacovigilance, further work is needed to improve its reliability and usability.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.013 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.004 | 0.002 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".