Phylogenetic prioritization of HIV-1 transmission clusters with viral lineage-level diversification rates
Bibliographic record
Abstract
Background and objectives: Public health officials faced with a large number of transmission clusters require a rapid, scalable and unbiased way to prioritize distribution of limited resources to maximize benefits. We hypothesize that transmission cluster prioritization based on phylogenetically derived lineage-level diversification rates will perform as well as or better than commonly used growth-based prioritization measures, without need for historical data or subjective interpretation. Methodology: 9822 HIV pol sequences collected during routine drug resistance genotyping were used alongside simulated sequence data to infer sets of phylogenetic transmission clusters via patristic distance threshold. Prioritized clusters inferred from empirical data were compared to those prioritized by the current public health protocols. Prioritization of simulated clusters was evaluated based on correlation of a given prioritization measure with future cluster growth, as well as the number of direct downstream transmissions from cluster members. Results: Empirical data suggest diversification rate-based measures perform comparably to growth-based measures in recreating public heath prioritization choices. However, unbiased simulated data reveals phylogenetic diversification rate-based measures perform better in predicting future cluster growth relative to growth-based measures, particularly long-term growth. Diversification rate-based measures also display advantages over growth-based measures in highlighting groups with greater future transmission events compared to random groups of the same size. Furthermore, diversification rate measures were notably more robust to effects of decreased sampling proportion. Conclusions and implications: Our findings indicate diversification rate-based measures frequently outperform growth-based measures in predicting future cluster growth and offer several additional advantages beneficial to optimizing the public health prioritization process.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".