Diagnostic Potential of Cross-Questionnaire Analysis for Depression and Memory Disorders Using Machine Learning Techniques
Bibliographic record
Abstract
Introduction Research shows a strong correlation between depression and memory disorders, suggesting the potential for cross-questionnaire data use in automated diagnostic systems. This study explores whether the Prospective and Retrospective Memory Questionnaire (PRMQ) can identify depressive symptoms and if the ZUNG Self-Rating Depression Scale (SDS) can predict memory-related disorders. Objectives To evaluate the effectiveness of using questionnaires intended for one mental disorder to diagnose another through machine learning models on data from a large-scale self-assessment online questionaire. Methods The study is part of the Memory and Depression Study: MANDY, conducted by the 1st Department of Psychiatry and the Department of Neurology of Papageorgiou Hospital, Thessaloniki, Greece. Data from 3340 participants were collected via an online survey containing the PRMQ, SDS, demographic data, and health-related questions. Four predictive tasks were designed: two for predicting depression using memory responses (D-from-M score and class) and two for predicting memory disorders using depression responses (M-from-D score and class). Machine learning models including LightGBM, AdaBoost, Support Vector Machines, and Logistic Regression were evaluated. Performance metrics included precision, recall, F1-score, and AUC-ROC (Figure 1). Results The LightGBM classifier was the top-performing model for the D-from-M class prediction task, achieving a precision of 0.75563, recall of 0.79125, an F1-score of 0.77303, and an AUC-ROC of 0.79319 on the test set (Table 1). This indicates a strong predictive capability for diagnosing depression from memory-related responses. The AdaBoost classifier had similar performance but was slightly inferior to LightGBM. For the M-from-D class task, the class imbalance (memory disorder prevalence at 5%) was a significant challenge. The best model, a Support Vector Classifier with ADASYN resampling, achieved a precision of 0.6, recall of 0.375, an F1-score of 0.46154, and an AUC-ROC of 0.86218. However, its performance was notably lower than LightGBM in predicting depression. Image 1: Image 2: Conclusions The PRMQ, combined with specific demographic and health-related questions, showed promise in predicting depression, with the LightGBM classifier as the best overall model. This underscores the potential for cross-questionnaire data utilization for diagnosing depression. Conversely, predicting memory disorders using the SDS was less effective, indicating the need for more targeted diagnostic tools. Future research should include neurocognitive and biomarker data to enhance diagnostic accuracy for memory-related conditions. Disclosure of Interest None Declared
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".