MétaCan
Menu
Back to cohort
Record W4392293824 · doi:10.1101/2024.02.28.580462

Phylogenetic signal in primate tooth enamel proteins and its relevance for paleoproteomics

2024· preprint· en· W4392293824 on OpenAlexaff
Ricardo Fong-Zazueta, Johanna Krueger, David M. Alba, Xènia Aymerich, Robin M. D. Beck, Enrico Cappellini, Guillermo Carrillo Martín, Omar Cirilli, Nathan L Clark, Omar E. Cornejo, Kyle Kai‐How Farh, Luis Ferrández-Peral, David Juan, Joanna L. Kelley, Lukas F. K. Kuderna, Jordan Little, Joseph D. Orkin, Ryan Paterson, Harvinder Pawar, Tomàs Marquès‐Bonet, Esther Lizano

Bibliographic record

VenuebioRxiv (Cold Spring Harbor Laboratory) · 2024
Typepreprint
Languageen
FieldMedicine
TopicBone and Dental Protein Studies
Canadian institutionsUniversité de Montréal
FundersAgencia Estatal de InvestigaciónFundación Bancaria Caixa d'Estalvis i Pensions de BarcelonaAgència de Gestió d'Ajuts Universitaris i de RecercaGeneralitat de CatalunyaEuropean CommissionSight Research UKMinisterio de Ciencia e InnovaciónNatural Environment Research CouncilCentres de Recerca de Catalunya
KeywordsPhylogenetic treeBiologyEvolutionary biologyCladePhylogeneticsProteomePhylogenetic networkComputational phylogeneticsComputational biologyGeneticsGene

Abstract

fetched live from OpenAlex

Abstract Ancient tooth enamel, and to some extent dentin and bone, contain characteristic peptides that persist for long periods of time. In particular, peptides from the enamel proteome (enamelome) have been used to reconstruct the phylogenetic relationships of fossil specimens and to estimate divergence times. However, the enamelome is based on only about 10 genes, whose protein products undergo fragmentation post mortem . Moreover, some of the enamelome genes are paralogous or may coevolve. This raises the question as to whether the enamelome provides enough information for reliable phylogenetic inference. We address these considerations on a selection of enamel-associated proteins that has been computationally predicted from genomic data from 232 primate species. We created multiple sequence alignments (MSAs) for each protein and estimated the evolutionary rate for each site and examined which sites overlap with the parts of the protein sequences that are typically isolated from fossils. Based on this, we simulated ancient data with different degrees of sequence fragmentation, followed by phylogenetic analysis. We compared these trees to a reference species tree. Up to a degree of fragmentation that is similar to that of fossil samples from 1-2 million years ago, the phylogenetic placements of most nodes at family level are consistent with the reference species tree. We found that the composition of the proteome influences the phylogenetic placement of Tarsiiformes. For the inference of molecular phylogenies based on paleoproteomic data, we recommend characterizing the evolution of the proteomes from the closest extant relatives to maximize the reliability of phylogenetic inference.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.002
Threshold uncertainty score0.010

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.004
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0010.001
Open science0.0000.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.018
GPT teacher head0.250
Teacher spread0.232 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuebioRxiv (Cold Spring Harbor Laboratory)Same topicBone and Dental Protein StudiesFrench-language works237,207