Open science datasets from PREVENT‐AD, a longitudinal cohort of pre‐symptomatic Alzheimer’s disease
Notice bibliographique
Résumé
Abstract Background In the last 10 years, the PResymptomatic EValuation of Experimental or Novel Treatments for Alzheimer Disease (PREVENT‐AD) research program has collected both biological and behavioral data longitudinally, using well‐recognized biomarkers, to better understand and prevent Alzheimer’s disease (AD). Our cohort is composed of older individuals with a family history of AD, but who were themselves cognitively unimpaired when enrolled in the study (aged 63 years old ± 5, at entry). To move AD research further, we recently shared this rich dataset with the global research community. Method We created a data sharing model including two levels of access based on data sensitivity and risk of potential re‐identification, framed with a series of usage terms. Preparation of the datasets included selection of the variables to share, quality controls, outlier analysis, de‐identification, and creation of detailed data dictionaries for proper re‐usability. Each research participant was retrospectively re‐consented. Result More than 90% of participants (349 out of 386) agreed to openly share their data. Repositories are accessible openly at https://openpreventad.loris.ca, to qualified researchers at https://registeredpreventad.loris.ca, and through the unified interface of the Canadian Open Neuroscience Platform. The shared data collected from 2012 to 2020 (Table 1) includes longitudinal multimodal magnetic resonance imaging, gene variants, neuropsychological assessments, neurosensory evaluations, subjective cognitive decline information, as well as multiple behavioral data (e.g. sleep quality, neuropsychiatric factors, personality traits, etc). Amyloid and tau measurements from cerebrospinal fluid (n=106; longitudinal) and positron emission tomography (n = 130) are also available on subsamples of participants. To date, more than 250 users have accessed the PREVENT‐AD datasets. Conclusion Creation of open datasets from sensitive human data requires technical, human, and financial investments. The challenge remains important as we must constantly deploy our sharing initiative efforts as new data is collected, and new modalities incorporated in the cohort. By offering this evolving resource to the research community, we aim to acknowledge the implication of our research participants by expanding the potential of the data they generously provided, generate a higher rate of new discoveries in AD pathogenesis, and contribute to the international efforts towards prevention of AD.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,008 | 0,046 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,003 | 0,006 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,003 | 0,004 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».