Using Twitter to Surveil the Opioid Epidemic in North Carolina: An Exploratory Study
Bibliographic record
Abstract
BACKGROUND: Over the last two decades, deaths associated with opioids have escalated in number and geographic spread, impacting more and more individuals, families, and communities. Reflecting on the shifting nature of the opioid overdose crisis, Dasgupta, Beletsky, and Ciccarone offer a triphasic framework to explain that opioid overdose deaths (OODs) shifted from prescription opioids for pain (beginning in 2000), to heroin (2010 to 2015), and then to synthetic opioids (beginning in 2013). Given the rapidly shifting nature of OODs, timelier surveillance data are critical to inform strategies that combat the opioid crisis. Using easily accessible and near real-time social media data to improve public health surveillance efforts related to the opioid crisis is a promising area of research. OBJECTIVE: This study explored the potential of using Twitter data to monitor the opioid epidemic. Specifically, this study investigated the extent to which the content of opioid-related tweets corresponds with the triphasic nature of the opioid crisis and correlates with OODs in North Carolina between 2009 and 2017. METHODS: Opioid-related Twitter posts were obtained using Crimson Hexagon, and were classified as relating to prescription opioids, heroin, and synthetic opioids using natural language processing. This process resulted in a corpus of 100,777 posts consisting of tweets, retweets, mentions, and replies. Using a random sample of 10,000 posts from the corpus, we identified opioid-related terms by analyzing word frequency for each year. OODs were obtained from the Multiple Cause of Death database from the Centers for Disease Control and Prevention Wide-ranging Online Data for Epidemiologic Research (CDC WONDER). Least squares regression and Granger tests compared patterns of opioid-related posts with OODs. RESULTS: The pattern of tweets related to prescription opioids, heroin, and synthetic opioids resembled the triphasic nature of OODs. For prescription opioids, tweet counts and OODs were statistically unrelated. Tweets mentioning heroin and synthetic opioids were significantly associated with heroin OODs and synthetic OODs in the same year (P=.01 and P<.001, respectively), as well as in the following year (P=.03 and P=.01, respectively). Moreover, heroin tweets in a given year predicted heroin deaths better than lagged heroin OODs alone (P=.03). CONCLUSIONS: Findings support using Twitter data as a timely indicator of opioid overdose mortality, especially for heroin.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".