Discriminating between schizophrenia subtypes using clustering and supervised learning
Bibliographic record
Abstract
Schizophrenia is a complex neuropsychiatric disorder affecting roughly 1% of the population, making it a significant public health concern [Sawa and Snyder, 2002].Presentation of schizophrenia includes positive (hallucinations, delusions) and negative (lack of motivation, inability to feel pleasure) symptoms, in addition to cognitive impairments [Owen et al., 2016].However, both symptom burden and associated brain alterations are highly heterogeneous and intimately linked to prognosis.To characterize the clinical heterogeneities, there is a need to develop methods that can predict symptom burden at the individual level.To serve this purpose, machine learning algorithms were implemented to first derive clinical subgroups from high-dimensional interrelated clinical information, and then subject-level classification was performed based on magnetic resonance imaging (MRI) derived neuroanatomical measures.Data from the SchizConnect website (http://schizconnect.org/) was used in this analysis, consisting of 104 patients and 63 normal controls.Unsupervised hierarchical clustering was performed on the symptom severity data of the patients.The 3-cluster solutions was chosen, following a stability analysis, representing patients with (1) high loads of both negative and positive symptoms, (2) predominantly positive symptoms, or (3) mild-symptomatology as compared to the average population.Demographic variables and the average cortical thickness in 78 brain regions defined by the Automated Anatomical Labeling atlas parcellation were used as input features This thesis is organized in the following manner:Chapter gives a brief background on schizophrenia, MRI, abnormal neuroanatomy in schizophrenia, machine learning, and data-driven methods for subtyping and prediction.Chapter provides the research statement and a breakdown of specific objectives.Chapter describes the dataset, MRI pre-processing, cortical thickness estimation, data-driven subtype definition, single-subject prediction, and classifier performance evaluation.Chapter outlines the results of the clustering and machine learning.Chapter provides a discussion of the major findings, limitations, and potential future directions.Chapter provides a summary of the thesis and gives concluding statements.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.000 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".