Identifying clinical clusters with distinct trajectories in first-episode psychosis through an unsupervised machine learning technique
Bibliographic record
Abstract
The extreme variability in symptom presentation reveals that individuals diagnosed with a first-episode psychosis (FEP) may encompass different sub-populations with potentially different illness courses and, hence, different treatment needs. Previous studies have shown that sociodemographic and family environment factors are associated with more unfavorable symptom trajectories. The aim of this study was to examine the dimensional structure of symptoms and to identify individuals' trajectories at early stage of illness and potential risk factors associated with poor outcomes at follow-up in non-affective FEP. One hundred and forty-four non-affective FEP patients were assessed at baseline and at 2-year follow-up. A Principal component analysis has been conducted to identify dimensions, then an unsupervised machine learning technique (fuzzy clustering) was performed to identify clinical subgroups of patients. Six symptom factors were extracted (positive, negative, depressive, anxiety, disorganization and somatic/cognitive). Three distinct clinical clusters were determined at baseline: mild; negative and moderate; and positive and severe symptoms, and five at follow-up: minimal; mild; moderate; negative and depressive; and severe symptoms. Receiving a low-dose antipsychotic, having a more severe depressive symptomatology and a positive family history for psychiatric disorders were risk factors for poor recovery, whilst having a high cognitive reserve and better premorbid adjustment may confer a better prognosis. The current study provided a better understanding of the heterogeneous profile of FEP. Early identification of patients who could likely present poor outcomes may be an initial step for the development of targeted interventions to improve illness trajectories and preserve psychosocial functioning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".