Understanding recruitment: outcomes associated with alternate methods for seed selection in respondent driven sampling
Bibliographic record
Abstract
BACKGROUND: Respondent driven sampling (RDS) was designed for sampling "hidden" populations and intended as a means of generating unbiased population estimates. Its widespread use has been accompanied by increasing scrutiny as researchers attempt to understand the extent to which the population estimates produced by RDS are, in fact, generalizable to the actual population of interest. In this study we compare two different methods of seed selection to determine whether this may influence recruitment and RDS measures. METHODS: Two seed groups were established. One group was selected as per a standard RDS approach of study staff purposefully selecting a small number of individuals to initiate recruitment chains. The second group consisted of individuals self-presenting to study staff during the time of data collection. Recruitment was allowed to unfold from each group and RDS estimates were compared between the groups. A comparison of variables associated with HIV was also completed. RESULTS: Three analytic groups were used for the majority of the analyses-RDS recruits originating from study staff-selected seeds (n = 196); self-presenting seeds (n = 118); and recruits of self-presenting seeds (n = 264). Multinomial logistic regression demonstrated significant differences between the three groups across six of ten sociodemographic and risk behaviours examined. Examination of homophily values also revealed differences in recruitment from the two seed groups (e.g. in one arm of the study sex workers and solvent users tended not to recruit others like themselves, while the opposite was true in the second arm of the study). RDS estimates of population proportions were also different between the two recruitment arms; in some cases corresponding confidence intervals between the two recruitment arms did not overlap. Further differences were revealed when comparisons of HIV prevalence were carried out. CONCLUSIONS: RDS is a cost-effective tool for data collection, however, seed selection has the potential to influence which subgroups within a population are accessed. Our findings indicate that using multiple methods for seed selection may improve access to hidden populations. Our results further highlight the need for a greater understanding of RDS to ensure appropriate, accurate and representative estimates of a population can be obtained from an RDS sample.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.050 | 0.203 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".