Comparison of Finite and Infinite Mixture Models for Capturing Compositional Heterogeneity Across Sites
Bibliographic record
Abstract
Phylogenetic modelling of the variation of the evolutionary process across sites from multispecies sequence alignments has garnered increasing attention over the last few decades.One of the main approaches, sometimes known as random effects modelling, adopts the view that the heterogeneity across observations is a result of the data set having been emitted from several different models, each drawn from a distribution.When little is known about the form of the across-site heterogeneity, finite mixture models provide discretizations of the unknown distribution into a pre-determined set of sub-models, or components.Choosing a level of discretization that is sufficiently fine-meshed to reflect the underlying heterogeneity is typically done from a set of likelihood-based model comparisons using different numbers of components.In the infinite mixture framework, accounting for the uncertainty regarding the number of components is another layer built into the model formulation (i.e., a hierarchical modelling framework), providing a rich non-parametric fitting of the distribution of acrosssite heterogeneity.Here, we use Bayesian cross-validation to compare a wide range of finite mixture models, along with the infinite mixture modelling approach known as categories, 'CAT', and gamma-distributed rates-across-sites approach.We study the model comparison approach on simulations, and apply it to five real multi-gene alignments.Our findings indicate that the potential improvement in model-fit from finite mixture models is attained when the number of components of the mixture is between 20 and 60.The magnitude of improvement from the mixture model is highly dependant on whether or not the gammadistributed rates-across-sites approach is invoked.Moreover, in all cases that we considered, the fit of the CAT-GTR+Γ model matched or exceeded the best-fitting finite mixture model.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.023 | 0.048 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.005 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".