The impact of horizontal gene transfer in shaping operons and protein interaction networks – direct evidence of preferential attachment
Bibliographic record
Abstract
BACKGROUND: Despite the prevalence of horizontal gene transfer (HGT) in bacteria, to this date there were few studies on HGT in the context of gene expression, operons and protein-protein interactions. Using the recently available data set on the E. coli protein-protein interaction network, we sought to explore the impact of HGT on genome structure and protein networks. RESULTS: We classified the E. coli genes into three categories based on their evolutionary conservation: a set of 2158 Core genes that are shared by all E. coli strains, a set of 1044 Non-core genes that are strain-specific, and a set of 1053 genes that were putatively acquired by horizontal transfer. We observed a clear correlation between gene expressivity (measured by Codon Adaptation Index), evolutionary rates, and node connectivity between these categories of genes. Specifically, we found the Core genes are the most highly expressed and the most slowly evolving, while the HGT genes are expressed at the lowest level and evolve at the highest rate. Core genes are the most likely and HGT genes are the least likely to be member of the operons. In addition, we found the Core genes on average are more highly connected than Non-core and HGT genes in the protein interaction network, however the HGT genes displayed a significantly higher mean node degree than the Core and Non-core genes in the defence COG functional category. Interestingly, HGT genes are more likely to be connected to Core genes than expected by chance, which suggest a model of differential attachment in the expansion of cellular networks. CONCLUSION: Results from our analysis shed light on the mode and mechanism of the integration of horizontally transferred genes into operons and protein interaction networks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".