The utility of visual appearance in predicting the composition of street opioids
Bibliographic record
Abstract
Background With the emergence of unregulated fentanyl, people who use unregulated opioids are increasingly relying on appearance in an effort to ascertain the presence of fentanyl and level of drug potency. However, the utility of visual inspection to identify drug composition in the fentanyl era has not been assessed. Methods We assessed client expectation, appearance, and composition of street drug samples being presented for drug checking. Results of a visual screening test were compared to fentanyl immunoassay strip testing. We calculated sensitivity, specificity and likelihood ratios (LR) to assess the accuracy of the common assumption that samples with a “pebbles” appearance contain fentanyl. Results In total, of the 2502 unregulated opioid samples tested, 1820 (73.5%) appeared as “pebbles”, of which 1729 (95.0%) tested positive for fentanyl for a sensitivity of 75.9% (95% Confidence Interval [CI]: 74.2–77.6) and specificity of 59.4% (95%CI: 57.5–61.3). Although, the odds of samples containing fentanyl was 4.60 (95%CI: 3.47−6.11) times higher among pebbles samples compared to non-pebble samples, the positive LR for pebbles to contain fentanyl was only 1.87 (CI: 1.59–2.19). The negative LR was more useful at 0.41 (95% CI: 0.36−0.46). Conclusions A positive screening test for pebbles is not strongly enough associated to be used as a proxy for detecting fentanyl. While the absence of the appearance of pebbles does somewhat reduce the likelihood of fentanyl being present in a given sample, the high prevalence of fentanyl and fentanyl analogues in the drug supply and the risks of consumption are such that public health providers should routinely advise people who use unregulated opioids against solely relying on visual characteristics of drugs as a harm reduction strategy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".