Phylogenetic Incongruence among Oncogenic Genital Alpha Human Papillomaviruses
Bibliographic record
Abstract
The human papillomaviruses (HPVs) have long been thought to follow a monophyletic pattern of evolution with little if any evidence for recombination between genomes. On the basis of this model, both oncogenicity and tissue tropism appear to have evolved once. Still, no systematic statistical analyses have shown whether monophyly is the rule across all HPV open reading frames (ORFs). We conducted a taxonomic analysis of 59 mucosal/genital HPVs using whole-genome and sliding-window similarity measures; maximum-parsimony, neighbor-joining, and Bayesian phylogenetic analyses; and localized incongruence length difference (LILD) analyses. The algorithm for the LILD analyses localized incongruence by calculating the tree length differences between constrained and unconstrained nodes in a total-evidence tree across all HPV ORFs. The process allows statistical evaluation of every ORF/node pair in the total-evidence tree. The most significant incongruence was observed at the putative high-risk (i.e., cancer-associated) node, the common oncogenic ancestor for alpha HPV species 9 (e.g., HPV type 16 [HPV16]), 11, 7 (e.g., HPV18), 5, and 6. Although these groups share early-gene homology, including high degrees of similarity among E6 and E7, groups 9 and 11 diverge from groups 7, 5, and 6 with respect to L2 and L1. The HPV species groups primarily associated with cervical and anogenital cancers appear to follow two distinct evolutionary paths, one conferred by the early genes and another by the late genes. The incongruence in the genital HPV phylogeny could have occurred from an early recombination event, an ecological niche change, and/or asymmetric genome convergence driven by intense selection. These data indicate that the phylogeny of the oncogenic HPVs is complex and that their evolution may not be monophyletic across all genes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".