Simulating within host human immunodeficiency virus 1 genome evolution in the persistent reservoir
Bibliographic record
Abstract
Abstract The complexities of viral evolution can be difficult to elucidate. Software simulating viral evolution provides powerful tools for exploring hypotheses of viral systems, especially in situations where thorough empirical data are difficult to obtain or parameters of interest are difficult to measure. Human immunodeficiency virus 1 (HIV-1) infection has no durable cure; this is primarily due to the virus’ ability to integrate into the genome of host cells, where it can remain in a transcriptionally latent state. An effective cure strategy must eliminate every copy of HIV-1 in this ‘persistent reservoir’ because proviruses can reactivate, even decades later, to resume an active infection. However, many features of the persistent reservoir remain unclear, including the temporal dynamics of HIV-1 integration frequency and the longevity of the resulting reservoir. Thus, sophisticated analyses are required to measure these features and determine their temporal dynamics. Here, we present software that is an extension of SANTA-SIM to include multiple compartments of viral populations. We used the resulting software to create a model of HIV-1 within host evolution that incorporates the persistent HIV-1 reservoir. This model is composed of two compartments, an active compartment and a latent compartment. With this model, we compared five different date estimation methods (Closest Sequence, Clade, Linear Regression, Least Squares, and Maximum Likelihood) to recover the integration dates of genomes in our model’s HIV-1 reservoir. We found that the Least Squares method performed the best with the highest concordance (0.80) between real and estimated dates and the lowest absolute error (all pairwise t tests: P < 0.01). Our software is a useful tool for validating bioinformatics software and understanding the dynamics of the persistent HIV-1 reservoir.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".