Using Named Entity Recognition to Identify Substances Used in the Self-medication of Opioid Withdrawal: Natural Language Processing Study of Reddit Data
Bibliographic record
Abstract
BACKGROUND: The cessation of opioid use can cause withdrawal symptoms. People often continue opioid misuse to avoid these symptoms. Many people who use opioids self-treat withdrawal symptoms with a range of substances. Little is known about the substances that people use or their effects. OBJECTIVE: The aim of this study is to validate a methodology for identifying the substances used to treat symptoms of opioid withdrawal by a community of people who use opioids on the social media site Reddit. METHODS: We developed a named entity recognition model to extract substances and effects from nearly 4 million comments from the r/opiates and r/OpiatesRecovery subreddits. To identify effects that are symptoms of opioid withdrawal and substances that are potential remedies for these symptoms, we deduplicated substances and effects by using clustering and manual review, then built a network of substance and effect co-occurrence. For each of the 16 effects identified as symptoms of opioid withdrawal in the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition, we identified the 10 most strongly associated substances. We classified these pairs as follows: substance is a Food and Drug Administration-approved or commonly used treatment for the symptom, substance is not often used to treat the symptom but could be potentially useful given its pharmacological profile, substance is a home or natural remedy for the symptom, substance can cause the symptom, or other or unclear. We developed the Withdrawal Remedy Explorer application to facilitate the further exploration of the data. RESULTS: scores of 92.1 (substances) and 81.7 (effects) on hold-out data. We identified 458 unique substances and 235 unique effects. Of the 130 potential remedies strongly associated with withdrawal symptoms, 54 (41.5%) were Food and Drug Administration-approved or commonly used treatments for the symptom, 17 (13.1%) were not often used to treat the symptom but could be potentially useful given their pharmacological profile, 13 (10%) were natural or home remedies, 7 (5.4%) were causes of the symptom, and 39 (30%) were other or unclear. We identified both potentially promising remedies (eg, gabapentin for body aches) and potentially common but harmful remedies (eg, antihistamines for restless leg syndrome). CONCLUSIONS: Many of the withdrawal remedies discussed by Reddit users are either clinically proven or potentially useful. These results suggest that this methodology is a valid way to study the self-treatment behavior of a web-based community of people who use opioids. Our Withdrawal Remedy Explorer application provides a platform for using these data for pharmacovigilance, the identification of new treatments, and the better understanding of the needs of people undergoing opioid withdrawal. Furthermore, this approach could be applied to many other disease states for which people self-manage their symptoms and discuss their experiences on the web.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".