Identification of metabolomics-based biomarker discovery in individuals with down syndrome utilizing kernel-tree model-enhanced explainable artificial intelligence methodology
Bibliographic record
Abstract
OBJECTIVE: This study aims to develop an explainable artificial intelligence (XAI) model integrated with machine learning (ML) to comprehensively investigate metabolic differences between individuals with Down syndrome (T21) and healthy controls (D21) and to identify novel/pathway-specific biomarkers. In this study, ML classifiers including AdaBoost, LightGBM, Random Forest, KTBoost, and XGBoost are applied to metabolomics data obtained from metabolomic analyses by high-resolution liquid chromatography-mass spectrometry (LC-MS) using blood plasma samples of 316 T21 and 103 D21 individuals, and the importance of metabolites is evaluated by XAI-based SHAP analysis. The KTBoost model shows the highest classification performance with an accuracy of 90.4% and area under the curve (AUC) of 95.9%, outperforming AdaBoost, LightGBM, Random Forest, and XGBoost. Significant downregulation and upregulation of some metabolites were observed in the T21 group compared to the D21 group. Metabolites such as vitamin C, taurolithocholic acid, sphingosine, and prostaglandin A2/B2/J2 are observed at low levels in the T21 group. In contrast, metabolites such as thymidine, tau-roursodeoxycholic acid, serine, and nervonic acid are elevated. SHAP analysis revealed that L-Citrulline, Kynurenin, Prostaglandin A2/B2/J2, Urate, and Pantothenate metabolites could be novel/pathway-specific biomarkers to differentiate the T21 group. This study revealed significant metabolic alterations in individuals with T21 and demonstrated the effectiveness of the combination of ML and XAI methods to identify novel/pathway-specific biomarkers. The findings may contribute to a better understanding of Down syndrome's molecular mechanisms and the development of future diagnostic and therapeutic strategies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".