Longitudinal Phylogenetic Surveillance Identifies Distinct Patterns of Cluster Dynamics
Bibliographic record
Abstract
OBJECTIVE: Through the application of simple, accessible, molecular epidemiology tools, we aimed to resolve the phylogenetic relationships that best predicted patterns of cluster growth using longitudinal population level drug resistance genotype data. METHODS: Analysis was performed on 971 specimens from drug naïve, first time HIV positive subjects collected in British Columbia between 2002 and 2005. A 1240bp fragment of the pol gene was amplified and sequenced with relationships among subtype B sequences inferred using Neighbour-Joining analysis. Apparent clusters of infections having both a mean within group distance <0.031 and bootstrap value >80% were systematically identified. The entire 2002-2005 dataset was then re-analyze to evaluate the relationship of subsequent infections to those identified in 2002. BED testing was used to identify recent infections (<156 days). RESULTS: Among the 2002 infections, 136 of 300 sequences sorted into 52 clusters ranging in size from 2 to 9 members. Aboriginal ethnicity and intravenous drug use were correlated, and both were linked to cluster membership in 2002. Although cluster growth between 2002 and 2005 was correlated with the size of the original cluster, more related infections were found in clusters seeded from nonclustered infections. Finally, all large growth clusters were seeded from infections that were much more likely to be recent. CONCLUSIONS: This population level phylogenetic analysis suggests that a greater increase in cluster size is associated with recently infected individuals, which may represent the leading edge of the epidemic. The most impressive increase in cluster size is seen originating from initially nonclustered infections. In contrast, smaller existing clusters likely describe historical patterns of transmission and do not substantially contribute to the ongoing epidemic. Application of this method for cross-sectional analysis of existing sequences from defined geographic regions may be useful in predicting trends in HIV transmission.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".