Mental Illness Diagnostic Criteria Can Be Simplified With Higher Symptom Prevalence and Correlations: A Simulation Study
Bibliographic record
Abstract
Introduction The diagnostic criteria of mental illnesses have been found to assign excessive weights to certain input symptoms, and several input symptoms may not be significantly associated with their diagnosis. This study aims to investigate whether we can use input symptoms assigned much weight to diagnose mental illnesses with similar diagnostic accuracy as using all input symptoms. Methods The symptoms of three mental conditions were simulated based on a published study: major depressive episodes, dysthymic disorder, and manic episodes. We simulated symptoms with 0.05, 0.1, 0.3, 0.5, or 0.7 prevalence and 0, 0.1, 0.4, 0.7, or 0.9 correlations. For each of the 25 combinations of symptom prevalence and correlations, we simulated 100,000 subjects, and diagnoses were made based on the Diagnostic and Statistical Manual of Mental Disorders, 4th edition, text revision. For each simulation, we used a decision tree model that used a symptom with the best diagnostic accuracy to separate a population into diseased and non-diseased groups. This model continued using other symptoms to further separate the groups into subgroups. This process was repeated until the diagnostic accuracy could not be improved based on cross-validation errors. All analyses were implemented with R (v4.2.3; R Development Core Team, Vienna, Austria) and RStudio (v2023.6.0.421; RStudio Team, Boston, MA). Results The diagnoses of major depressive episodes, dysthymic disorder, and manic episodes required 15, 11, and 14 symptoms, respectively. There were opportunities to use fewer symptoms to approximate the diagnoses with 92% or higher sensitivities and specificities with certain combinations of symptom prevalence and correlations. For major depressive episodes, using two symptoms ("Depressed mood" and "Loss of interest or pleasure in daily activities" for more than two weeks) in the major criteria could diagnose the condition, with 100% sensitivity and specificity in some circumstances. Occasionally, the diagnosis of dysthymic disorder might be used to approximate the diagnosis of major depressive episodes. Conclusion There may lie opportunities to screen or follow up the diagnosis of major depressive episodes, dysthymic disorder, and manic episodes using fewer input symptoms, with at least 92% sensitivities and specificities. These opportunities exist in various combinations of symptom prevalence and correlations. However, there is a lack of real-world data on psychiatric symptoms and interventions to take advantage of these opportunities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.010 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".