Testing Phylogenetic Placement Accuracy of <scp>DNA</scp> Barcode Sequences on a Fish Backbone Tree: Implications of Backbone Tree Completeness and Species Representation
Bibliographic record
Abstract
Advancements in DNA sequencing technology have facilitated the generation of a vast number of DNA sequences, posing opportunities and challenges for constructing large phylogenetic trees. DNA barcode sequences, particularly COI, represent extensive orthologous sequences suitable for phylogenetic analysis. Phylogenetic placement analysis offers a promising method to integrate COI data into tree-building efforts, yet the impacts of backbone tree completeness and species composition remain under-explored. Using a dataset comprising 27 genes and 4520 species of bony fishes, we assessed the accuracy of phylogenetic inference by "placing" COI sequences onto backbone trees. The backbone tree completeness was varied by subsampling 20%, 40%, 60%, 80%, and 99% of the total species separately, followed by placement of those missing species based on their COI sequences using software packages EPA-ng and APPLES. We also compared the effects of biased, random, and stratified sampling strategies; the latter ensured the representation of all major lineages (Family) of bony fish. Our findings indicate that the placement accuracy is consistently high across all levels of backbone tree completeness, where 70%-78% missing species are correctly placed (by EPA-ng) in the same locations as the reference tree derived from the complete data. High completeness produces slightly high placement accuracy, although in many cases the differences are nonsignificant. For example, at the 99% completeness level with stratified sampling, EPA-ng placed 78% missing species correctly, and when only considering placement with high confidence (LWR > 0.9), the percentage is 87%. Additionally, stratified sampling outperforms random sampling in most cases, and biased sampling has the worst performance. The likelihood-based EPA-ng consistently provide higher accurate placements than the distance-based APPLES. In conclusion, COI-based placement analysis represents a potential route of using the available vast barcoding data for building large phylogenetic trees.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".