The silent majority: The typical Canadian sex worker may not be who we think
Bibliographic record
Abstract
BACKGROUND: Most sex worker population studies measure population at discrete points in time and very few studies have been done in industrialized democracies. The purpose of this study is to consider how time affects the population dynamics of contact sex workers in Canada using publicly available internet advertising data collected over multiple years. METHODS: 3.6 million web pages were collected from advertising sites used by contact sex workers between November, 2014 and December, 2016 inclusive. Contacts were extracted from ads and used to identify advertisers. First names were used to estimate the number of workers represented by an advertiser. Counts of advertisers and names were adjusted for missing data and overcounting. Two approaches for correcting overcounts are compared. Population estimates were generated weekly, monthly and for the two year period. The length of time advertisers were active was also estimated. Estimates are also compared with related research. RESULTS: Canadian sex workers typically advertised individually or in small collectives (median name count 1, IQR 1-2, average 1.8, SD 4.4). Advertisers were active for a mean of 73.3 days (SD 151.8, median 14, IQR 1-58). Advertisers were at least 83.5% female. Respectively the scaled weekly, monthly, and biannual estimates for female sex workers represented 0.2%, 0.3% and 2% of the 2016 Canadian female 20-49 population. White advertisers were the most predominant ethnic group (53%). CONCLUSIONS: Sex work in Canada is a more pervasive phenomenon than indicated by spot estimates and the length of the data collection period is an important variable. Non-random samples used in qualitative research in Canada likely do not reflect the larger sex worker population represented in advertising. The overall brevity of advertising activity suggests that workers typically exercise agency, reflecting the findings of other Canadian research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.007 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.010 | 0.005 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.010 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".