Principal component analysis as an efficient method for capturing multivariate brain signatures of complex disorders—<scp>ENIGMA</scp> study in people with bipolar disorders and obesity
Bibliographic record
Abstract
Multivariate techniques better fit the anatomy of complex neuropsychiatric disorders which are characterized not by alterations in a single region, but rather by variations across distributed brain networks. Here, we used principal component analysis (PCA) to identify patterns of covariance across brain regions and relate them to clinical and demographic variables in a large generalizable dataset of individuals with bipolar disorders and controls. We then compared performance of PCA and clustering on identical sample to identify which methodology was better in capturing links between brain and clinical measures. Using data from the ENIGMA-BD working group, we investigated T1-weighted structural MRI data from 2436 participants with BD and healthy controls, and applied PCA to cortical thickness and surface area measures. We then studied the association of principal components with clinical and demographic variables using mixed regression models. We compared the PCA model with our prior clustering analyses of the same data and also tested it in a replication sample of 327 participants with BD or schizophrenia and healthy controls. The first principal component, which indexed a greater cortical thickness across all 68 cortical regions, was negatively associated with BD, BMI, antipsychotic medications, and age and was positively associated with Li treatment. PCA demonstrated superior goodness of fit to clustering when predicting diagnosis and BMI. Moreover, applying the PCA model to the replication sample yielded significant differences in cortical thickness between healthy controls and individuals with BD or schizophrenia. Cortical thickness in the same widespread regional network as determined by PCA was negatively associated with different clinical and demographic variables, including diagnosis, age, BMI, and treatment with antipsychotic medications or lithium. PCA outperformed clustering and provided an easy-to-use and interpret method to study multivariate associations between brain structure and system-level variables. PRACTITIONER POINTS: In this study of 2770 Individuals, we confirmed that cortical thickness in widespread regional networks as determined by principal component analysis (PCA) was negatively associated with relevant clinical and demographic variables, including diagnosis, age, BMI, and treatment with antipsychotic medications or lithium. Significant associations of many different system-level variables with the same brain network suggest a lack of one-to-one mapping of individual clinical and demographic factors to specific patterns of brain changes. PCA outperformed clustering analysis in the same data set when predicting group or BMI, providing a superior method for studying multivariate associations between brain structure and system-level variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".