Zero-inflated Bayesian factor analysis model with skew-normal priors for modeling microbiome data
Bibliographic record
Abstract
Abstract Background Advancements in next-generation sequencing have transformed our understanding of host-microbe interactions, revealing links between microbial composition and chronic conditions such as diabetes, Crohn’s disease, and others. However, the analysis of microbiome data is complex due to its high-dimension as well as additional statistical characteristics. One primary objective is to achieve effective dimensionality reduction while simultaneously accounting for the data’s compositional nature and zero inflation. Existing probabilistic models are often based on the assumption that the log-ratio-transformed compositions are normally distributed. This assumption is problematic because it can fail to capture the significant skewness inherent in these transformed compositions. Results We propose a new model called the Zero-Inflated Factor Analysis Logistic Skew-Normal Multinomial (ZIFA-LSNM) model. ZIFA-LSNM integrates a zero-inflation component to handle excess zeros, employs factor analysis for dimensionality reduction, and, critically, utilizes skew-normal priors on the latent factors to explicitly model data asymmetry. Posterior inference is performed using an efficient variational inference algorithm. Through simulation studies and real data analysis, the ZIFA-LSNM model is shown to have improved performance in parameter recovery and composition estimation compared to its Gaussian-based counterparts. Conclusion The ZIFA-LSNM model demonstrates that explicitly accounting for skewness in the latent factor structure can substantially improve inference in commonly-observed microbiome data. The proposed model, therefore, offers a flexible and scalable framework allowing for the improved analysis of the complex relationships between microbial communities and human health.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.026 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".