MétaCan
Menu
Back to cohort
Record W6907291636 · doi:10.20381/ruor-31296

Automated Detection of Substance Use Through Social Mining and its Prediction Ability in the Canadian Population

2025· dissertation· en· W6907291636 on OpenAlexaboutno aff

Bibliographic record

VenueUniversity of Ottawa - Library · 2025
Typedissertation
Languageen
FieldPsychology
TopicMental Health via Writing
Canadian institutionsnot available
Fundersnot available
KeywordsSubstance usePsychological interventionSocial mediaPopulationSubstance abuseConsumption (sociology)Process (computing)

Abstract

fetched live from OpenAlex

According to the latest WHO report in 2024, there is a significant increase in substance use disorders and environmental harms around the world. The report highlights that alcohol consumption was responsible for 2.6 million deaths annually, representing 4.7% of all world deaths, while psychoactive drug use accounted for 0.6 million deaths. The number of drug users increased to 292 million in 2022, reflecting a 20% rise over 10 years. Automated detection of different substance uses through social media can be an effective and practical observational tool for the global substance use problem. Automated detection of online communication has multiple applications, including helping people at-risk and protecting them by predicting and monitoring the early signs of risks on time. Our system can be used by individuals with authority (such as parents or doctors) to detect and monitor different substance users. It could raise an alarm to the relevant individuals to take necessary interventions for the early signs of substance use associated with the flagged posts. This thesis describes the process for classifying online posts to detect substance use problems as early as possible. We began by utilizing two datasets of annotated social media posts to train several classification models that predict whether these posts indicate signs of substance use. We assessed the performance of several traditional and recent deep learning models. Different CNN-based, RNN-based, BERT-based, and GPT models were found to be promising approaches in detecting substance users from their posts. GPT-4o, using a few-shot learning model, outperformed other models with 89.44% F1-score. Also, we built different user-level detection models for common substances (cannabis and alcohol). For cannabis user detection, GPT-4o using a few-shot learning model was the best-performing model with 85.22% F1-score, while the DeBERTa-v3 model was the best-performing model with 65.50% F1-score for alcohol user detection. As a second objective, these models were used for the automated detection of different substance use at the population level in Canada. A common practice for substance use detection at the population level involves conducting surveys via phone calls or interviews; however, this approach is both time-consuming and expensive. Understanding Canadian trends in alcohol and drug use is crucial for developing and evaluating effective policies and programs at both the national and provincial levels. Examining social media posts can serve as a flexible alternative for identifying several substance use problems across Canada. We detected the population-level use of cannabis and alcohol from 2015 to 2018, based on representative samples. Then, we compared these results of the same years' official statistics from Health Canada for the two substances. We used the estimated reports from Health Canada until 2019. Given the lack of annotated data for several substances, such as alcohol, we proposed a data augmentation technique that increased the information within the training phase by building several artificial training sets. Then, we applied the best generalized model (as mentioned before) for population-level detection. The results for population-level detection for both cannabis and alcohol were promising for the tested years and comparable with the results of the Health Canada surveys. The cannabis user detection achieved a difference of 5% or less from the governmental estimations for the nine Canadian provinces included in this study. Similarly, the alcohol user detection achieved a difference of 6.5% or less for the same group of provinces under study. To the best of our knowledge, this is the first study to propose the detection of substance use through social media for an entire country.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.480
Threshold uncertainty score0.831

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.001
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.283
Teacher spread0.253 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueUniversity of Ottawa - LibrarySame topicMental Health via WritingFrench-language works237,207