Network Distribution and Respondent-Driven Sampling (RDS) Inference About People Who Inject Drugs in Ottawa, Ontario
Bibliographic record
Abstract
Respondent-driven sampling (RDS) is very useful in collecting data from individuals in hidden populations, where a sampling frame does not exist. It starts with researchers choosing initial respondents from a group which may be involved in taboo or illegal activities, after which they recruit other peers who belong to the same group. Analysis results in unbiased estimates of population proportions though with strong assumptions about the underlying social network and RDS recruitment process. These assumptions bear little resemblance to reality, and thus compromise the estimation of any means, population proportions or variances inferred from studies. The topology of the contact network, denoted by the number of links each person has, provides insight into the processes of infectious disease spread. The overall objective of the thesis is to identify the topology of an injection drug use network, and critically review the methods developed to produce estimates. The topology of people who inject drugs (PWID) collected by RDS in Ottawa, 2006 was compared with a Poisson distribution, an exponential distribution, a power-law distribution, and a lognormal distribution. The contact distribution was then evaluated against a small-world network characterized by high clustering and low average distances between individuals. Last a systematic review of the methods used to produce RDS mean and variance estimates was conducted. The Poisson distribution, a type of random distribution, was not an appropriate fit for PWID network. However, the PWID network can be classified as a small world network organised with many connections and short distances between people. Prevention of transmission in such networks should be focussed on the most active people (clustered individuals and hubs) as intervention with any others is less effective. The systematic review contained 32 articles which included the development and evaluation of 12 RDS mean and 6 variance estimators. Overall, the majority of estimators perform roughly the same, with the exception of RDSIEGO which outperformed the 6 other RDS mean estimators. The Tree bootstrap variance estimate does not rely on modelling RDS as a first order Markov (FOM) process, which seems to be the main limitation of the other existing estimators. The lack of FOM as an assumption and the flexible application of this variance estimator to any RDS point estimate make the Tree bootstrapping estimator a more efficient choice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.101 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.006 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".