<scp>RepeatOBserver</scp> : Tandem Repeat Visualisation and Putative Centromere Detection
Bibliographic record
Abstract
Tandem repeats play an important role in centromere structure, subtelomeric regions, DNA methylation, recombination and the regulation of gene activity. Analysis of their distribution in genomes offers a potential means for predicting putative centromere locations, which continues to be a challenge for genome annotation. Here we present RepeatOBserver (https://github.com/celphin/RepeatOBserverV1), a new tool for visualising repeat patterns and identifying putative centromere locations, using a Fourier transform of DNA walks. RepeatOBserver can identify and visualise a broad range of perfect and imperfect repeats (3-5000 bp long) in genome assemblies without any a priori knowledge of repeat sequences or the need for optimising parameters. RepeatOBserver heatmaps can distinguish between tandem and retrotransposon repeats. We analysed 159 chromosomes with experimentally-verified centromere positions from 12 plant and animal species. We find that 93% of experimentally-verified tandem repeat centromeres occur in regions of low sequence diversity and 97% of retrotransposon centromeres occur in regions with a high abundance of repeat lengths. Depending on the centromere type predicted by the heatmaps, putative centromere locations can be predicted using either a genomic Shannon diversity index or a repeat abundance sum. RepeatOBserver can also locate other regions of interest including potential neocentromeres and gene copy variation. Split and inverted tandem repeats at inversion boundaries suggest that chromosomal inversions or mis-assemblies can also be located. RepeatOBserver is a flexible tool for comprehensive characterisation of repeat patterns that can be used to visualise and identify a variety of regions of interest in genome assemblies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".