2019 Undergraduate Big Data Challenge: Big Data of Recreational Drugs
Bibliographic record
Abstract
This paper aims to determine if the legalization of recreational cannabis in Colorado and Washington has had an impact on the trends of opioid overdose deaths in these states.Datasets were collected from organizations including the National Survey on Drug Use and Health (NSDUH) and the Centers for Disease Control and Prevention (CDC).The central target of analysis from these datasets is the number of opioid overdose deaths prior to and after the year of recreational cannabis legalization.This analysis is performed using linear and quadratic regression models, comparing the projections of the number of opioid overdose deaths made prior to the year of legalization with the actual number of opioid overdose deaths following legalization.Linear regression models were primarily used with the exception of cases in which a quadratic regression model represented the data more accurately.A confidence interval of 95% was used for the model projection.Through these methods, the authors found that there was no significant correlation between opioid overdose deaths in Colorado and Washington and the legalization of recreational cannabis.While the actual data of opioid overdose deaths did trend downward in most cases following cannabis legalization, it did not decrease to such an extent that it could not be explained by an error in the model: the data did not fall outside of the confidence interval.The downward trend of the actual data appears to closely follow the previously existing downward trend and varies little from the projections made before the legalization of recreational cannabis.Although the actual data displays a downward trend, the models suggest that the trend is upwards overall.Despite the lack of a strong correlation, it may be too recent to draw a definite conclusion.As more data is collected and more locations legalize cannabis for recreational use, revisiting the topic may yield different conclusions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.069 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.006 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.005 | 0.005 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.003 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".