Artificial intelligence-based framework for Alzheimer’s disease diagnosis via video vision transformer
Bibliographic record
Abstract
Objective: Alzheimer's disease (AD) is a progressive neurodegenerative disorder that leads to cognitive decline and memory impairment, posing a public health concern in aging populations. Early and accurate detection of AD using non-invasive imaging biomarkers remains a critical clinical need for timely intervention and disease management. This study aims to develop an advanced artificial intelligence (AI)-based diagnostic framework, ViTranZheimer, that leverages video vision transformers to analyze magnetic resonance imaging (MRI) and improve the accuracy of AD classification. Methods: This study presents 'ViTranZheimer,' an AD diagnosis approach that leverages video transformers to analyze MRI volumes. Our proposed deep learning framework aims to improve the accuracy and sensitivity of AD diagnosis, equipping clinicians with a tool for early detection and intervention. We exploit the temporal dependencies between slices by treating the MRI volumes as videos to capture intricate structural relationships. We evaluated ViTranZheimer on the publicly available Alzheimer's Disease Neuroimaging Initiative (ADNI): Complete 3Yr 3T data collection, which includes 351 T1-weighted MRI scans categorized into normal controls (NC = 129), mild cognitive impairment (MCI = 145), and AD = 77 groups. Each MRI volume was preprocessed using spatial normalization and skull stripping, then modeled as a video sequence for input to a Video Vision Transformer (ViViT). The model was trained from scratch using 10-fold stratified cross-validation and optimized with the Adam optimizer over 500 epochs. Classification performance was evaluated using accuracy, precision, recall, F1-score, and area under the ROC curve (AUC). Statistical comparison was conducted using the Wilcoxon signed-rank test against two baseline models: a convolutional neural network with bidirectional long short-term memory (CNN-BiLSTM), and a vision transformer with bidirectional long short-term memory (ViT-BiLSTM). Results: < 0.05). Conclusion: ViTranZheimer demonstrates strong potential for accurate and early Alzheimer's disease diagnosis using non-invasive MRI data. By leveraging video vision transformers, the model provides a promising tool for clinical decision support in neurodegenerative disease detection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".