Estimating the Population Size of People Who Inject Drugs in 3 Cities in Zambia: Capture-Recapture, Successive Sampling, and Bayesian Consensus Estimation Methods
Bibliographic record
Abstract
Background: Accurate population size estimates (PSE) of key populations-those disproportionately affected by HIV-are critical to forecast need and inform HIV prevention and treatment programs, though they can be difficult to ascertain due to low visibility of these groups. In Zambia, reliable estimates on the number of people who inject drugs are limited, inhibiting public health response. Objective: We sought to estimate the population size of people who inject drugs in 3 large cities in Zambia, assess how PSEs vary across different estimation methods, and explore the strengths and limitations of each approach. Methods: We applied 2-source capture-recapture (2S-CRC), 3-source capture-recapture (3S-CRC), and successive sampling population size estimation (SS-PSE) methods in Lusaka, Livingstone, and Ndola, Zambia. 3S-CRC methods included location-based 2S-CRC in combination with a respondent-driven sampling (RDS) survey. Data were collected from November 2021 to February 2022 and analyzed using a Bayesian nonparametric latent class model. SS-PSEs were produced using the RDS recruitment and network sizes. Kruskal tests and general linear models were used to examine sociodemographic and behavioral factors associated with being captured in 2S-CRC among RDS participants. Final city population estimates, incorporating 3S-CRC and SS-PSE with imputed visibility estimates, were generated using a Bayesian consensus estimator. Results: Bayesian consensus PSEs ranged between 0.5% and 1.8% of the adult male population and were below 1% of the total adult population in each city. Consensus estimates were highest in Lusaka (3700, 95% credible interval [CRI] 1500-7500), followed by Ndola (2200, 95% CRI 1600-2900) and Livingstone (1200, 95% CRI 900-1,900). There was variability in estimates by method, with SS-PSE with imputed visibility generally providing the lowest estimates across cities, excluding Lusaka. Across methods, PSEs and uncertainty bounds (95% confidence interval [CI] or CRI depending on method) ranged from 1510 (95% CRI 1030-2070) to 4350 (95% CI 1410-18,890) in Lusaka, 360 (95% CI 290-530) to 2620 (95% CRI 1510-4680) in Livingstone, and 760 (95% CI 390-3060) to 4030 (95% CRI 960-5480) in Ndola. In all cities, fewer recaptures occurred in capture 3 (RDS) than with location sampling via 2S-CRC. Though results varied across cities, RDS participants captured through 2S-CRC differed from those captured solely through RDS in sociodemographic and behavioral risk factors, including housing, education, injection or needle sharing frequency, time since last injection, receipt of drug treatment, and experience with a peer educator in at least one city. Conclusions: This study used rigorous methods to produce PSEs in Zambia, and is the first to produce these for major geographies in the country. Through RDS, 3S-CRC reached people who inject drugs with distinct characteristics that were less accessible via location-based sampling (2S-CRC), yielding a PSE that may better reflect the population and informing the Bayesian consensus estimate. Findings from this study can guide program planning and future surveillance activities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".