Examination of HIV-1 diversity and evolution by a bioinformatics approach
Bibliographic record
Abstract
HIV-1 genetic diversity is a major obstacle for developing an effective vaccine. My hypothesis is that HIV-1 genetic diversity can be characterized and that cross-clade immunogens can be predicted at the population level. I systematically investigated positive selection (PS) pressures on HIV-1 Env and Gag proteins based on the analysis of the sequences collected from the Los Alamos Sequence Database. I identified PS sites, investigated PS patterns, correlated PS with the known functional sites of the two proteins, calculated frequencies of HLA alleles targeting CTL epitopes, and compared PS patterns among major subtypes. The results showed that PS pressure was widely dispersed across the entire regions of both HIV-1 Env and Gag proteins, suggesting the conserved regions are under host immune response pressure. The neutralizing antibody, non-neutralizing antibody, and CTL responses were found to be the major forces driving genetic diversity of HIV-1 env and gag genes at population level. However, PS pressures on both Env and Gag proteins remain stable over time, suggesting genetic diversity of HIV-1 driven by host immune responses changed very little over the last 29 years. Furthermore, the results also demonstrated that up to 70% PS sites were shared among the major HIV-1 clades, implying the existence of cross-clade immunogenicity. A number of potential cross-clades immunogens were predicted to elicit CTL or neutralizing antibody responses from Env and Gag proteins. I also detected a significant correlation between HLA allele frequencies and host CTL responses elicited by Accessory/Regulator’s proteins at population level. Moreover, I detected an association between the frequency of HLA-B7 supertype and the number of identified optimal CTL epitopes. The results suggest HLA class I allele frequencies in a population influence the evolution of HIV-1. I also systematically evaluated the utility of ultra-deep pyrosequencing to characterize genetic diversity of HIV-1 gag genes within quasispecies. The results showed that ultra-deep pyrosequencing of amplified HIV genes is a better method than the traditional Sanger-clone-based method in the comprehensive characterization of genetic diversity of HIV-1 quasispecies, especially in detecting low frequency variations. In conclusion, my thesis provides important information for rational design of an effective HIV-1 vaccine.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".