Predicting superagers: a machine learning approach utilizing gut microbiome features
Bibliographic record
Abstract
Objective: Cognitive decline is often considered an inevitable aspect of aging; however, recent research has identified a subset of older adults known as "superagers" who maintain cognitive abilities comparable to those of younger individuals. Investigating the neurobiological characteristics associated with superior cognitive function in superagers is essential for understanding "successful aging." Evidence suggests that the gut microbiome plays a key role in brain function, forming a bidirectional communication network known as the microbiome-gut-brain axis. Alterations in the gut microbiome have been linked to cognitive aging markers such as oxidative stress and inflammation. This study aims to investigate the unique patterns of the gut microbiome in superagers and to develop machine learning-based predictive models to differentiate superagers from typical agers. Methods: We recruited 161 cognitively unimpaired, community-dwelling volunteers aged 60 years or from dementia prevention centers in Seoul, South Korea. After applying inclusion and exclusion criteria, 115 participants were included in the study. Following the removal of microbiome data outliers, 102 participants, comprising 57 superagers and 45 typical agers, were finally analyzed. Superagers were defined based on memory performance at or above average normative values of middle-aged adults. Gut microbiome data were collected from stool samples, and microbial DNA was extracted and sequenced. Relative abundances of bacterial genera were used as features for model development. We employed the LightGBM algorithm to build predictive models and utilized SHAP analysis for feature importance and interpretability. Results: The predictive model achieved an AUC of 0.832 and accuracy of 0.764 in the training dataset, and an AUC of 0.861 and accuracy of 0.762 in the test dataset. Significant microbiome features for distinguishing superagers included Alistipes, PAC001137_g, PAC001138_g, Leuconostoc, and PAC001115_g. SHAP analysis revealed that higher abundances of certain genera, such as PAC001138_g and PAC001115_g, positively influenced the likelihood of being classified as superagers. Conclusion: Our findings demonstrate the machine learning-based predictive models using gut-microbiome features can differentiate superagers from typical agers with a reasonable performance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".