Modeling Site-and-Branch-Heterogeneity with GFmix
Notice bibliographique
Résumé
Phylogenetic trees are often inferred from protein sequences sampled from diverse taxa across the tree of life. The compositions of these amino acid sequences may be heterogeneous across both sites and branches, particularly if deep phylogenetic divergences are the focus. Under some conditions, failure to model this compositional heterogeneity can lead to phylogenetic artefacts. However, the computational cost of phylogenetic inference with models accounting for compositional heterogeneity can be prohibitive. The originally proposed site-and-branch-heterogeneous GFmix model accounts for changing relative frequencies of G, A, R, and P (GARP) vs. F, Y, M, I, N, and K (FYMINK) amino acids resulting from extreme variation in G+C content among taxa. This GFmix model modifies a fitted site-heterogeneous profile mixture model in a branch-specific manner using parameters that reflect branch-specific amino acid compositions. This approach has been shown to improve likelihoods and reduce compositional artefacts. However, the original implementation of the model includes constraints which may sacrifice accuracy for computability and is limited to modeling variation in GARP/FYMINK composition. Here we investigate the properties of the original GFmix model in greater depth and present several improvements to the model. The improved GFmix models permit fewer constraints on branch-specific composition parameters, allow modeling of user-defined compositional heterogeneity, and provide for full maximum-likelihood optimization of parameters. We have also developed new methods for detecting compositional heterogeneity directly from sequence data. Analyses of simulated site-and-branch-heterogeneous data indicates that the improved GFmix models better estimate branch-specific compositions and branch lengths in heterogeneous trees. We applied the various versions of the GFmix model to a real dataset with known compositional heterogeneity artefacts. We find that the most complex GFmix model with full maximum likelihood parameter optimization consistently supports the correct tree over the artefactual tree with improved likelihoods. All implementations of the GFmix model and related scripts are available from https://www.mathstat.dal.ca/~tsusko/software.html.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».