MétaCan
Menu
Back to cohort
Record W3107095885 · doi:10.1038/s41598-020-77220-w

A comparison of machine learning methods for survival analysis of high-dimensional clinical data for dementia prediction

2020· article· en· W3107095885 on OpenAlexfundno aff
Annette Spooner, Emily Chen, Arcot Sowmya, Perminder S. Sachdev, Nicole A. Kochan, Julian N. Trollor, Henry Brodaty

Bibliographic record

VenueScientific Reports · 2020
Typearticle
Languageen
FieldMedicine
TopicDementia and Cognitive Impairment Research
Canadian institutionsnot available
FundersNational Institute on AgingNational Institute of Biomedical Imaging and BioengineeringCanadian Institutes of Health ResearchNational Institutes of HealthGenentechIXICOH. Lundbeck A/SServierEisaiNorthern California Institute for Research and EducationPfizerBiogenBioClinicaF. Hoffmann-La RocheUniversity of Southern CaliforniaMeso Scale DiagnosticsMedical Research CouncilEli Lilly and CompanyU.S. Department of DefenseAlzheimer's Disease Neuroimaging InitiativeNovartis Pharmaceuticals CorporationBristol-Myers SquibbAlzheimer's AssociationNational Health and Medical Research CouncilFoundation for the National Institutes of Health
KeywordsDementiaFeature selectionComputer scienceConcordanceMachine learningArtificial intelligenceClinical trialCohortNeuroimagingMissing dataDiseaseData miningMedicinePsychiatryInternal medicine

Abstract

fetched live from OpenAlex

Data collected from clinical trials and cohort studies, such as dementia studies, are often high-dimensional, censored, heterogeneous and contain missing information, presenting challenges to traditional statistical analysis. There is an urgent need for methods that can overcome these challenges to model this complex data. At present there is no cure for dementia and no treatment that can successfully change the course of the disease. Machine learning models that can predict the time until a patient develops dementia are important tools in helping understand dementia risks and can give more accurate results than traditional statistical methods when modelling high-dimensional, heterogeneous, clinical data. This work compares the performance and stability of ten machine learning algorithms, combined with eight feature selection methods, capable of performing survival analysis of high-dimensional, heterogeneous, clinical data. We developed models that predict survival to dementia using baseline data from two different studies. The Sydney Memory and Ageing Study (MAS) is a longitudinal cohort study of 1037 participants, aged 70-90 years, that aims to determine the effects of ageing on cognition. The Alzheimer's Disease Neuroimaging Initiative (ADNI) is a longitudinal study aimed at identifying biomarkers for the early detection and tracking of Alzheimer's disease. Using the concordance index as a measure of performance, our models achieve maximum performance values of 0.82 for MAS and 0.93 For ADNI.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.044
metaresearch head score (Gemma)0.075
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.044
Threshold uncertainty score0.233

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0440.075
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0060.003
Science and technology studies0.0010.001
Scholarly communication0.0020.003
Open science0.0010.001
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.209
GPT teacher head0.517
Teacher spread0.308 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations291
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueScientific ReportsSame topicDementia and Cognitive Impairment ResearchFrench-language works237,207