F38. PHENOTYPIC CLUSTERING OF BIPOLAR DISORDER SUPPORTS STRATIFICATION BY LITHIUM RESPONSIVENESS, NOT DIAGNOSTIC SUBTYPES
Bibliographic record
Abstract
Background Bipolar disorder (BD) affects 1% of the global population and functionally impairs 83% of the afflicted. Although BD is a highly heritable and biological disorder, phenotypic heterogeneity complicates elucidation of its etiology. While genetic studies have been fruitful, the effect sizes of genetic markers are generally too small to support clinical applications. Phenotypic stratification of BD may resolve this heterogeneity and provide more homogeneous groups for future studies. One common stratification approach involves categorizing BD into Types I and II, but it has been criticized for imposing categorical boundaries on a dimensional psychopathological condition, let alone the contested validity of BD-II as a diagnosis. A promising alternative subtype is responsiveness to prophylactic lithium. Clinical phenotypic profiles differentiate excellent from poor lithium responders, and lithium responsiveness is heritable, suggesting a genetic component. The present study adjudicates between stratification strategies based on (A) BD I/II subtype and (B) lithium responsiveness, by ascertaining which approach captures more phenotypic information. Specifically, we use detailed clinical data from lithium-treated patients with BD-I or II. Using their detailed clinical profiles, we identify data-driven phenotypic clusters, and measure the information these clusters carry about either BD subtype or lithium responsiveness. Our results may inform diagnostic nosology and stratification approaches for future genetic studies of BD. Methods We included adult patients with BD-I or II (N = 477 across four sites) who were treated with lithium as their principal mood stabilizer for at least one year. Treatment responsiveness was defined using the dichotomized Alda score. We performed hierarchical clustering on phenotypes defined by over 50 features, covering demographics, clinical course , family history, suicide behaviour , and comorbid conditions. We then measured the amount of information that inferred clusters carried about (A) BD subtype or (B) lithium responsiveness using adjusted mutual information (AMI) scores. Detailed phenotypic profiles across clusters were then evaluated with univariate comparisons. Results Two clusters were identified (n = 52 and n = 425), which captured significantly less information about BD subtype (AMI range 0.004 to 0.011 [SE range: 2e-4 to 4e-4]) than lithium responsiveness, (AMI 0.033 to 0.133 [1e-3 to 2e-3]). The smaller cluster had disproportionately more lithium responders (n = 42 [80.8%] vs.108 [25.4%]; p = 0.026), attention-deficit hyperactivity disorder (5 [9.6%] vs. 15 [3.5%]; p = 0.026), and learning disabilities (4 [7.7%] vs. 15 [3.5%]; p = 0.026). Discussion Detailed clinical phenotypes may offer more information about lithium responsiveness than diagnostic subtype, supporting lithium responsiveness rather than BD-I/II classification as a valid approach to stratification in clinical samples.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".