MétaCan
Menu
Back to cohort
Record W3133501805 · doi:10.1186/s13195-021-00879-4

Data analysis with Shapley values for automatic subject selection in Alzheimer’s disease data sets using interpretable machine learning

2021· article· en· W3133501805 on OpenAlexfundno aff
Louise Bloch, Christoph M. Friedrich

Bibliographic record

VenueAlzheimer s Research & Therapy · 2021
Typearticle
Languageen
FieldMedicine
TopicDementia and Cognitive Impairment Research
Canadian institutionsnot available
FundersNational Institute of Biomedical Imaging and BioengineeringCanadian Institutes of Health ResearchNational Institutes of HealthGenentechIXICOH. Lundbeck A/SServierEisaiNorthern California Institute for Research and EducationBioClinicaF. Hoffmann-La RocheUniversity of Southern CaliforniaBiogenU.S. Department of DefenseMeso Scale DiagnosticsAlzheimer's Disease Neuroimaging InitiativeNovartis Pharmaceuticals CorporationPfizerEli Lilly and CompanyBristol-Myers SquibbNational Institute on AgingAlzheimer's AssociationFoundation for the National Institutes of Health
KeywordsOverfittingArtificial intelligenceMachine learningFeature selectionRandom forestComputer scienceTest setCross-validationNeuroimagingData miningMedicineArtificial neural network

Abstract

fetched live from OpenAlex

BACKGROUND: For the recruitment and monitoring of subjects for therapy studies, it is important to predict whether mild cognitive impaired (MCI) subjects will prospectively develop Alzheimer's disease (AD). Machine learning (ML) is suitable to improve early AD prediction. The etiology of AD is heterogeneous, which leads to high variability in disease patterns. Further variability originates from multicentric study designs, varying acquisition protocols, and errors in the preprocessing of magnetic resonance imaging (MRI) scans. The high variability makes the differentiation between signal and noise difficult and may lead to overfitting. This article examines whether an automatic and fair data valuation method based on Shapley values can identify the most informative subjects to improve ML classification. METHODS: An ML workflow was developed and trained for a subset of the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort. The validation was executed for an independent ADNI test set and for the Australian Imaging, Biomarker and Lifestyle Flagship Study of Ageing (AIBL) cohort. The workflow included volumetric MRI feature extraction, feature selection, sample selection using Data Shapley, random forest (RF), and eXtreme Gradient Boosting (XGBoost) for model training as well as Kernel SHapley Additive exPlanations (SHAP) values for model interpretation. RESULTS: The RF models, which excluded 134 of the 467 training subjects based on their RF Data Shapley values, outperformed the base models that reached a mean accuracy of 62.64% by 5.76% (3.61 percentage points) for the independent ADNI test set. The XGBoost base models reached a mean accuracy of 60.00% for the AIBL data set. The exclusion of those 133 subjects with the smallest RF Data Shapley values could improve the classification accuracy by 2.98% (1.79 percentage points). The cutoff values were calculated using an independent validation set. CONCLUSION: The Data Shapley method was able to improve the mean accuracies for the test sets. The most informative subjects were associated with the number of ApolipoproteinE ε4 (ApoE ε4) alleles, cognitive test results, and volumetric MRI measurements.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.034
metaresearch head score (Gemma)0.094
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.034
Threshold uncertainty score0.182

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0340.094
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0030.002
Science and technology studies0.0010.002
Scholarly communication0.0030.002
Open science0.0010.002
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.269
GPT teacher head0.475
Teacher spread0.206 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations64
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueAlzheimer s Research & TherapySame topicDementia and Cognitive Impairment ResearchFrench-language works237,207