Patient Stratification Using Metabolomics to Address the Heterogeneity of Psychosis
Bibliographic record
Abstract
Abstract Psychosis is a symptomatic endpoint with many causes, complicating its pathophysiological characterization and treatment. Our study applies unsupervised clustering techniques to analyze metabolomic data, acquired using 2 different tandem mass spectrometry (MS-MS) methods, from an unselected group of 120 patients with psychosis. We performed an independent analysis of each of the 2 datasets generated, by both hierarchical clustering and k-means. This led to the identification of biochemically distinct groups of patients while reducing the potential biases from any single clustering method or datatype. Using our newly developed robust clustering method, which is based on patients consistently grouped together through different methods and datasets, a total of 20 clusters were ascertained and 78 patients (or 65% of the original cohort) were placed into these robust clusters. Medication exposure was not associated with cluster formation in our study. We highlighted metabolites that constitute nodes (cluster-specific metabolites) vs hubs (metabolites in a central, shared, pathway) for psychosis. For example, 4 recurring metabolites (spermine, C0, C2, and PC.aa.C38.6) were discovered to be significant in at least 8 clusters, which were identified by at least 3 different clustering approaches. Given these metabolites were affected across multiple biochemically different patient subgroups, they are expected to be important in the overall pathophysiology of psychosis. We demonstrate how knowledge about such hubs can lead to novel antipsychotic medications. Such pathways, and thus drug targets, would not have been possible to identify without patient stratification, as they are not shared by all patients, due to the heterogeneity of psychosis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".