Differential degeneration of the ACTAGT sequence among Salmonella: a reflection of distinct nucleotide amelioration patterns during bacterial divergence
Bibliographic record
Abstract
When bacteria diverge, they need to adapt to the new environments, such as new hosts or different tissues of the same host, by accumulating beneficial genomic variations, but a general scenario is unknown due to the lack of appropriate methods. Here we profiled the ACTAGT sequence and its degenerated forms (i.e., hexa-nucleotide sequences with one of the six nucleotides different from ACTAGT) in Salmonella to estimate the nucleotide amelioration processes of bacterial genomes. ACTAGT was mostly located in coding sequences but was also found in several intergenic regions, with its degenerated forms widely scattered throughout the bacterial genomes. We speculated that the distribution of ACTAGT and its degenerated forms might be lineage-specific as a consequence of different selection pressures imposed on ACTAGT at different genomic locations (in genes or intergenic regions) among different Salmonella lineages. To validate this speculation, we modelled the secondary structures of the ACTAGT-containing sequences conserved across Salmonella and many other enteric bacteria. Compared to ACTAGT at conserved regions, the degenerated forms were distributed throughout the bacterial genomes, with the degeneration patterns being highly similar among bacteria of the same phylogenetic lineage but radically different across different lineages. This finding demonstrates biased amelioration under distinct selection pressures among the bacteria and provides insights into genomic evolution during bacterial divergence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".