The Volume and Tone of Twitter Posts About Cannabis Use During Pregnancy: A Scoping Review Protocol
Bibliographic record
Abstract
Background: Cannabis use has increased in Canada since its legalization in 2018, includingamong pregnant women who may be motivated to use cannabis to reduce symptoms ofnausea and vomiting. However, a growing body of research suggests that cannabis useduring pregnancy may harm the developing fetus. As a result, patients increasingly seekmedical advice from online sources, but these platforms may also spread anecdotaldescriptions or misinformation. Given the possible disconnect between online messaging andevidence-based research about the effects of cannabis use during pregnancy, there is apotential for advice taken from social media to cause harm.Objectives: To quantify the volume and tone of English-language posts related to cannabisuse in pregnancy from January 2012 to July 2021.Methods: Modelling published frameworks for scoping reviews, we will collect publiclyavailable posts from Twitter that mention cannabis use during pregnancy and employ theTwitter Application Programming Interface (API) for Academic Research to extract data fromtweets, including public metrics such as the number of likes, retweets and quotes, as well ashealth effect mentions, sentiment, location and users interests. These data will be used toquantify how cannabis use during pregnancy is discussed on Twitter and to build a qualitativeprofile of supportive and opposing posters.Results: The CHEO Research Ethics Board reviewed our project and granted an exemptionin May 2021. As of September 2021, we have gained approval to use the Twitter API forAcademic Research and have developed a preliminary search strategy that returns over 2million unique tweets posted between 2012 and 2020.Conclusions: Understanding how Twitter is being used to discuss cannabis use duringpregnancy will help public health agencies and healthcare providers assess the messagingpatients may be receiving and develop communication strategies to counter misinformation,especially in geographical regions where legalization is recent or imminent. Most importantly,we foresee that our findings will assist expecting families in making informed choices aboutwhere they choose to access advice about using cannabis during pregnancy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.096 | 0.111 |
| Meta-epidemiology (narrow) | 0.003 | 0.005 |
| Meta-epidemiology (broad) | 0.008 | 0.009 |
| Bibliometrics | 0.020 | 0.012 |
| Science and technology studies | 0.005 | 0.004 |
| Scholarly communication | 0.007 | 0.007 |
| Open science | 0.004 | 0.008 |
| Research integrity | 0.006 | 0.005 |
| Insufficient payload (model declined to judge) | 0.058 | 0.013 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".