Diagnostic performance of HIV risk assessment tools for identifying pre-exposure prophylaxis candidates: a systematic review and meta-analysis
Bibliographic record
Abstract
Background: To support the implementation of HIV pre-exposure prophylaxis (PrEP), we conducted a systematic review and meta-analysis evaluating the diagnostic performance of HIV risk assessment tools in predicting HIV infection. Methods: We searched MEDLINE, Embase, and CINAHL for observational studies published between January 1, 1998, and May 13, 2024 that reported on the diagnostic performance of HIV risk assessment tools. We calculated pooled area under the curve (pAUC) values using inverse variance methods, with sensitivity and specificity reported at common cutoffs (PROSPERO registration number: CRD42024543975). Findings: Of 3704 publications, 27 met our criteria. Twelve studies on men who have sex with men (MSM) assessed nine tools, with four extensively validated, predominantly in U.S. populations. SexPro exhibited the highest performance (pAUC: 0.75), while HIRI-MSM (pAUC: 0.69), Menza (pAUC: 0.63), and SDET (pAUC: 0.66) demonstrated moderate predictive ability, with considerable heterogeneity. For cisgender women, twelve African studies evaluated six tools, with VOICE being the only extensively validated tool (pAUC: 0.65 for adult females; 0.62 for adolescent and young women). Although additional tools were available for subgroups within Africa, there were no tools for cisgender women outside Africa. Among other populations, DHRS demonstrated good discrimination for general U.S. adults (pAUC: 0.80), as did the HIV Prevalence Risk Score for African mixed populations (AUC: 0.70), Kahle for heterosexual serodiscordant couples in Africa (pAUC: 0.73), and ARCH-IDU for people who use drugs in the U.S. (pAUC: 0.72). Sensitivity and specificity varied by cutoffs. Tool items fell into six domains: sexual activities, substance use, clinical factors, demographics, reproductive health, and other factors, with complexity differing by population and context. Interpretation: Validated tools can help identify HIV risk in some populations, but tools are still needed to promote equitable PrEP access for subpopulations such as cisgender women outside Africa. Public health programs and clinicians should consider incorporating up-to-date, local data to enhance the relevance and effectiveness of existing tools. Funding: This work was supported by the Canadian Institutes of Health Research (Grant number PCS - 183410). DHST is supported by a Tier 2 Canada Research Chair in Biomedical HIV/STI Prevention.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.049 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.016 | 0.006 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".