Clustering patients by depression symptoms to predict venlafaxine ER antidepressant efficacy: Individual patient data analysis
Bibliographic record
Abstract
To identify clusters of patients with major depressive disorder (MDD) based on the baseline 17-item Hamilton Rating Scale for Depression (HAM-D17) items and to evaluate the efficacy of venlafaxine extended release (VEN) vs placebo, and the potential effect of dose on efficacy, in each cluster. Cluster analysis was performed to identify clusters based on standardized HAM-D17 item scores of individual patient data at baseline from 9 double-blind, placebo-controlled studies of VEN for MDD. Change from baseline in HAM-D17 total score was analyzed using a mixed-effects model for repeated measures for each cluster; response and remission rates at week 8 were analyzed using logistic regression. Discontinuation rates were also evaluated in each cluster. In 2599 patients, 3 patient clusters were identified, characterized as High modified Core (mCore) Symptoms/High Anxiety (cluster 1), High mCore Symptoms/Medium Anxiety (cluster 2), and Medium mCore Symptoms/Medium Anxiety (cluster 3). Significant effects of VEN vs placebo were observed on change from baseline in HAM-D17 total score at week 8 for both clusters 1 and 2 (both P < 0.001), but not for cluster 3. In cluster 3, a significant treatment effect of VEN was observed at week 8 in the lower-dose subgroup but not in the higher-dose subgroup. All-cause discontinuation rates were significantly higher in placebo than VEN in each cluster. Three unique clusters of patients were identified differing in baseline mCore symptoms and anxiety. Cluster membership may predict efficacy outcomes and contribute to dose effects in patients treated with VEN. NCT01441440; other studies included in this analysis were conducted before the requirement to register clinical studies took effect.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".