A cross-sectional study of explainable machine learning in Alzheimer’s disease: diagnostic classification using MR radiomic features
Bibliographic record
Abstract
Introduction: Alzheimer's disease (AD) even nowadays remains a complex neurodegenerative disease and its diagnosis relies mainly on cognitive tests which have many limitations. On the other hand, qualitative imaging will not provide an early diagnosis because the radiologist will perceive brain atrophy on a late disease stage. Therefore, the main objective of this study is to investigate the necessity of quantitative imaging in the assessment of AD by using machine learning (ML) methods. Nowadays, ML methods are used to address high dimensional data, integrate data from different sources, model the etiological and clinical heterogeneity, and discover new biomarkers in the assessment of AD. Methods: In this study radiomic features from both entorhinal cortex and hippocampus were extracted from 194 normal controls (NC), 284 mild cognitive impairment (MCI) and 130 AD subjects. Texture analysis evaluates statistical properties of the image intensities which might represent changes in MRI image pixel intensity due to the pathophysiology of a disease. Therefore, this quantitative method could detect smaller-scale changes of neurodegeneration. Then the radiomics signatures extracted by texture analysis and baseline neuropsychological scales, were used to build an XGBoost integrated model which has been trained and integrated. Results: The model was explained by using the Shapley values produced by the SHAP (SHapley Additive exPlanations) method. XGBoost produced a f1-score of 0.949, 0.818, and 0.810 between NC vs. AD, MC vs. MCI, and MCI vs. AD, respectively. Discussion: These directions have the potential to help to the earlier diagnosis and to a better manage of the disease progression and therefore, develop novel treatment strategies. This study clearly showed the importance of explainable ML approach in the assessment of AD.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".