A176 UTILITY OF MACHINE LEARNING FOR SERUM METABOLOMIC DATA ANALYSIS IN PEDIATRIC CROHN DISEASE
Bibliographic record
Abstract
Abstract Background The pathogenesis of pCD remains poorly understood, but evidence suggests roles for genetics, environment, immune response, and gut microbes. Microbial changes can contribute to chronic inflammation and correlate with disease severity. Metabolomics reflects interactions between host immune and gut microbial function by quantifying compounds in biological samples. Therefore, metabolomics provides a unique opportunity to gain insight into pCD pathogenesis. Aims To correlate disease severity, metabolites, and clinical data by applying machine learning algorithms in pediatric Crohn Disease (pCD). Methods ImageKids is a multicenter, prospective, cohort observational study, conducted to develop magnetic resonance enterography (MRE) indices for pCD. Paired serum specimens were collected at study initiation (Visit One; V1) and completion (Visit Four; V4; 18 months) for 120 pCD patients. Serum from patients with representative clinical scenarios and paired samples was analyzed at The Metabolomics Innovation Centre (TMIC; University of Alberta) and 131 metabolites were identified. Metabolites were analyzed via Unsupervised (U.ML) and Supervised (S.ML) Machine Learning algorithms based on Scikit-learn library in Python. Principal Component Analysis (PCA) was used to identify the variation pattern of the patients’ metabolome. Classifiers and regression algorithms were trained to assess correlation with disease activity. Results Results were available for the 56 paired samples. U.ML demonstrated distinct metabolome profiles with V1 clustering mainly attributed to aspartic acid, glutamic acid, and kynurenine. V4 clustering was mainly attributed to spermidine, spermine, total dimethylarginine. Furthermore, demographics was found as an important environmental factor driving distinct patterns of the metabolomics profile. After training different classifiers and regressors with S. ML algorithms, metabolome data were correlated with disease severity (defined by C-reactive protein and fecal calprotectin). Isoleucine, p-hydroxyhippuric acid, and putrescine were the top three compounds associated with disease severity. The accuracy of our classification models was of 80% and the coefficient of determination of our regression models was 0.5 Conclusions Metabolomic analysis can provide insight into disease pathogenesis and help predict disease severity among pCD patients. The correlation between metabolomics and disease severity might allow a better understanding of changes in host-microbe interactions and introduce new diagnostic or therapeutic options. Funding Agencies CIHR
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".