Discovery of the first Tn630 member and the closest homolog of IS630 from viruses
Bibliographic record
Abstract
IS630/Tc1/mariner (ITm) represents the most widely distributed superfamily of DNA transposons in nature. Currently, bioinformatics research on ITm members primarily involves collecting data of existing and emerging members and organizing them into new groups or families. In the present study, our survey revealed that Tc1 and IS630 members have a broad host range, spanning across all six biological kingdoms (bacteria, fungi, plantae, animalia, archaea and protista) and viruses. The primary discoveries include the first Tn630 member-Tn630-NC1 and the closest homolog of IS630 from viruses-Tc1-C#1. By incorporating our discoveries into existing knowledge, we proposed a model to elucidate the formation of composite transposons. Organization of Tc1 and IS630 members into groups across biological kingdoms facilitates data collection for future research, particularly on their horizontal transfer between different kingdoms. The formation of composite transposons may result from asymmetric of terminal inverted repeats. IS630 should be merged with Tc1 into a single family IS630/Tc1. Furthermore, IS630 and its homologs constitute a valuable resource for studying horizontal gene transfer between gut bacteria and phages, opening up new avenues for research in this field.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".