MétaCan
Menu
← Back to cohort

Testing the Effects of Speciation and Mutation Rates on Distance-Based Phylogenetic Tree Construction Accuracy Using EcoSim

2014· article· en· W2944795414 on OpenAlexfundno aff
Ryan Scott, Robin Gras

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Phylogenetic Studies
Canadian institutionsnot available
FundersNatural Sciences and Engineering Research Council of Canada
KeywordsPhylogenetic treeGenetic algorithmComputer scienceTree (set theory)Mutation rateMutationComputational phylogeneticsArtificial intelligenceBiologyEvolutionary biologyMathematicsPhylogenetic networkGeneticsCombinatoricsGene

Abstract

fetched live from OpenAlex

Biologists regularly construct phylogenetic trees during their research in order to understand the evolutionary history of their subjects.Typically the actual phylogenetic tree is not known, and thus the phylogenetic tree produced is only an estimate.There are two main categories of phylogenetic tree construction, distance-based methods and character-based methods.In all cases, the effects of mutation rate and genetic compactness of species clusters on phylogenetic tree construction accuracy are not well understood.In order to test the correctness of a particular method, it is imperative to perform the study in a system for which the actual phylogenetic tree is known.Thus, realistic and complex simulations in which evolution and speciation occur provide a perfect platform for such a study.EcoSim is such a simulation, and is a simulation in which predator and prey agents interact, evolve, and speciate.Agents in EcoSim possess a complex, evolving, heritable behavioral model which provides meaningful evolution.Here, EcoSim was used as a platform on which to test the effects of mutation rate and speciation threshold on the accuracy of several distance-based phylogenetic tree construction methods.EcoSim has the ability to record all speciation events during a run, therefore we were able to construct the actual phylogenetic trees allowing us to properly compare estimation methods.We created four EcoSim run types: mutation rate increased (MRI), mutation rate decreased (MRD), speciation threshold increased (STI), and speciation threshold decreased (STD).We created five runs of each type, as well as five runs using the standard EcoSim configuration that all lasted 10000 time-steps.At various time-steps throughout each run, we performed Neighbor-Joining, UPGMA, and Fitch-Margoliash tree construction (with and without bootstrapping) on random subsets of 10 species that existed during that time-step, and compared the results to the actual phylogenetic tree using the symmetric distance metric.We found that Neighbor-Joining and Fitch-Margoliash performed nearly equally well, whereas UPGMA performed relatively poorly overall.Further, we found that an increase in speciation rate leads to performance losses in phylogenetic tree construction whereas modifying the mutation rate typically leads to performance gains.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.005
metaresearch head score (Gemma)0.023
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.006
Threshold uncertainty score0.025

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0050.023
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.002
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.013
GPT teacher head0.237
Teacher spread0.224 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2014
Admission routes1
Has abstractyes

Explore more

Same topicGenomics and Phylogenetic Studies→French-language works237,207→