Effective Characterisation of the Complete Orang-Utan Mitochondrial DNA Control Region, in the Face of Persistent Focus in Many Taxa on Shorter Hypervariable Regions
Bibliographic record
Abstract
The hypervariable region I (HVRI) is persistently used to discern haplotypes, to distinguish geographic subpopulations, and to infer taxonomy in a range of organisms. Numerous studies have highlighted greater heterogeneity elsewhere in the mitochondrial DNA control region, however-particularly, in some species, in other understudied hypervariable regions. To assess the abundance and utility of such potential variations in orang-utans, we characterised 36 complete control-region haplotypes, of which 13 were of Sumatran and 23 of Bornean maternal ancestry, and compared polymorphisms within these and within shorter HVRI segments predominantly analysed in prior phylogenetic studies of Sumatran (~385 bp) and Bornean (~323 bp) orang-utans. We amplified the complete control region in a single PCR that proved successful even with highly degraded, non-invasive samples. By using species-specific primers to produce a single large amplicon (~1600 bp) comprising flanking coding regions, our method also serves to better avoid amplification of nuclear mitochondrial insertions (numts). We found the number, length and position of hypervariable regions is inconsistent between orang-utan species, and that prior definitions of the HVRI were haphazard. Polymorphisms occurring outside the predominantly analysed segments were phylogeographically informative in isolation, and could be used to assign haplotypes to comparable clades concordant with geographic subpopulations. The predominantly analysed segments could discern only up to 76% of all haplotypes, highlighting the forensic utility of complete control-region sequences. In the face of declining sequencing costs and our proven application to poor-quality DNA extracts, we see no reason to ever amplify only specific 'hypervariable regions' in any taxa, particularly as their lengths and positions are inconsistent and cannot be reliably defined-yet this strategy predominates widely. Given their greater utility and consistency, we instead advocate analysis of complete control-region sequences in future studies, where any shorter segment might otherwise have proven the region of choice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".