Machine learning using multimodal clinical, electroencephalographic, and magnetic resonance imaging data can predict incident depression in adults with epilepsy: A pilot study
Bibliographic record
Abstract
OBJECTIVE: This study was undertaken to develop a multimodal machine learning (ML) approach for predicting incident depression in adults with epilepsy. METHODS: We randomly selected 200 patients from the Calgary Comprehensive Epilepsy Program registry and linked their registry-based clinical data to their first-available clinical electroencephalogram (EEG) and magnetic resonance imaging (MRI) study. We excluded patients with a clinical or Neurological Disorders Depression Inventory for Epilepsy (NDDI-E)-based diagnosis of major depression at baseline. The NDDI-E was used to detect incident depression over a median of 2.4 years of follow-up (interquartile range [IQR] = 1.5-3.3 years). A ReliefF algorithm was applied to clinical as well as quantitative EEG and MRI parameters for feature selection. Six ML algorithms were trained and tested using stratified threefold cross-validation. Multiple metrics were used to assess model performances. RESULTS: Of 200 patients, 150 had EEG and MRI data of sufficient quality for ML, of whom 59 were excluded due to prevalent depression. Therefore, 91 patients (41 women) were included, with a median age of 29 (IQR = 22-44) years. A total of 42 features were selected by ReliefF, none of which was a quantitative MRI or EEG variable. All models had a sensitivity > 80%, and five of six had an F1 score ≥ .72. A multilayer perceptron model had the highest F1 score (median = .74, IQR = .71-.78) and sensitivity (84.3%). Median area under the receiver operating characteristic curve and normalized Matthews correlation coefficient were .70 (IQR = .64-.78) and .57 (IQR = .50-.65), respectively. SIGNIFICANCE: Multimodal ML using baseline features can predict incident depression in this population. Our pilot models demonstrated high accuracy for depression prediction. However, overall performance and calibration can be improved. This model has promise for identifying those at risk for incident depression during follow-up, although efforts to refine it in larger populations along with external validation are required.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".