Latent Structure and Causal Mediation Models for Human Microbiome Sequencing Data
Bibliographic record
Abstract
The exploration of human microbiome sequencing data has emerged as a popular research area in recent medical science. Nonetheless, owing to inherent characteristics of microbiome data such as excess zeros, high skewness, sparsity, and a profusion of outliers, the analytical process faces some substantial hurdles. Consequently, the overarching aim of this thesis is to develop statistical methodologies that enable comprehension of human microbiome sequencing data analysis, encompassing investigation of latent structures and detection of potentially mediating roles played by certain microbial taxa between exposure/treatment to the interest outcomes. The first objective is to develop clustering algorithms which can be employed for microbiome datasets, seeking to unveil underlying data structure and potential subcategories. A novel approach, employing a customizable distance metric, is developed to eliminate the presence-absence bias linked with sparse count data while accurately measuring sample distances.Furthermore, when analyzing the role of microbial taxa as mediators - linking exposure to a study outcome via causal mediation analysis of microbiomes - single or multiple microbial entities may emerge as mediators. My second objective is to construct a causal mediation framework utilizing a weighting-based model specifically designed for a single zero-inflated microbial, demonstrating that our proposed framework precisely quantifies the mediation effect of specific taxa with minimum bias. Lastly, managing high-dimensional zero-inflated microbiome mediators within a counterfactual framework becomes crucial given that these mediators are correlated and often present with zero inflation. My third objective is to develop a method that incorporates the aforementioned distance metric in Objective One, specifically employed for microbiome data, capable of solidly assessing the frequency of zeros in the dataset, therefore tackling the dimensionality reduction faced by the microbiome data to prevent overfitting during mediation analysis. All approaches are subjected to comprehensive theoretical deduction and simulation evaluation. Concurrently, I demonstrate the model performance through real data applications. A gut microbiome dataset of Parkinson's disease was used for the first objective. The implementation of the second and third objectives utilizes the DIABIMMUNE Study to examine the mediation effect of breastfeeding on allergic reactions of infants through their gut microbiome.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.032 | 0.070 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".