From algorithms to zero-shot learning: how artificial intelligence is redefining gerontological research
Notice bibliographique
Résumé
By 2050, the global population of adults (≥60 years) is expected to double, to 2.1 billion (World Health Organization, 2024). To support older adults’ independence and participation in society, communities will need to create age-friendly and inclusive healthcare systems that manage multiple chronic conditions, expand long-term and dementia care, and redesign housing and transportation to ensure accessibility (Chum et al., 2022; Rowe et al., 2016). The intersection of artificial intelligence (AI) and gerontological science is accelerating our understanding of common conditions and management of aging-related cognitive and behavioral changes in the aging population (Chen et al., 2023). AI has emerged as a transformative tool in gerontological research, enabling the analysis of large-scale health data to enhance diagnostics, predictive risk modeling, and overall care delivery for older adults (Bernal et al., 2024). AI algorithms hold promise in supporting early disease detection and diagnosis, identifying conditions like delirium, and predicting adverse events like falls through machine learning (O’Connor et al., 2022; Pagali et al., 2023). Additionally, AI enables continuous monitoring via wearable sensors and smart home systems, allowing for tailored interventions and improved resource allocation in healthcare. Researchers are also leveraging AI to analyze longitudinal aging datasets, such as medical imaging and electronic health records, to map trajectories of cognitive decline and other age-related changes, offering new insights into dementia progression and functional aging patterns (Kourtis et al., 2019). This special issue of the Psychological Sciences section of The Journals of Gerontology, Series B, explores how metrics, including digital biomarkers and phenotypes, can improve early detection, disease monitoring, and intervention in aging populations. It highlights the innovative use of AI techniques ranging from traditional machine learning (ML) and deep learning (DL) to natural language processing and computer vision in advancing the precision measurement of cognitive, psychological, and social functioning among older adults. Studies included in this special issue fall into two broad categories. The first category comprises investigations focused on the use of AI/ML to analyze or improve the classification and prediction of aging and disease-related characteristics based on traditional, yet often high-dimensional, data types (e.g., demographic, clinical, and neuropsychological data). These analyses demonstrate how AI/ML can be leveraged to extract new insights from existing data, where traditional approaches have been less well-suited for data-driven analyses. For example, Jiang & Yu (2025) used ML to identify factors related to social resilience in older adults across three datasets, thereby supporting the prioritization of interventions to improve the social adaptability of older adults. Other papers focus on using AI/ML to improve the prediction of health and Alzheimer’s disease and related dementia (AD/ADRD)-related characteristics and outcomes based on existing data. Petersen et al. (2025) developed risk scores for predicting amyloid and tau biomarkers using ML (AutoScore) applied to data collected as part of the Alzheimer’s Disease Neuroimaging dataset. This approach enabled hierarchical combinations of demographic, neuropsychological, genetic, and imaging characteristics for prediction, in a manner that could enhance participant screening for clinical trials. Other contributions leveraged computer vision algorithms to better capture features of traditional neuropsychological task data. Hu et al. (2025) used deep learning neural network models to improve dementia classification from clock drawing images collected as part of the National Health and Aging Trends Study. In addition to improving efficiency, this approach allowed the development of continuous scores that were better able to balance sensitivity and specificity. Other studies in this special issue also highlight how AI/ML can be applied to prediction based on time-series data. Using transformer-based models, Weiss et al. (2025) achieved a twofold improvement in (next-wave) mortality prediction using data from the Health and Retirement Study collected over 29 years. In addition to advances in prediction, this special issue includes work validating approaches to ML-based classification. Nguyen & Chang (2025) compared results from different data-driven clustering algorithms (nonnegative matrix factorization and model-based clustering) with distinct heuristics, applied to demographic, clinical, neuropsychological, and brain morphometric data. Their analysis supports the comparability of these clustering algorithms for generating convergent insights. The second category of contribution included in this special issue involves the application of AI/ML to data generated from current and next-generation digital technologies, including passive data collected from sensors and smartphone-based digital cognitive assessments. These data are often temporally dense and may not be amenable to many traditional data analytic approaches; however, they can be used to generate novel digital phenotypes that capture essential information about health and behavior. Oravecz et al. (2025) utilized high-frequency longitudinal data from smartphone-based cognitive assessments to identify computational cognitive markers of long-term forgetting, learning, and within-person variability related to cognitive aging and mild cognitive impairment. Importantly, these markers reflect changes in cognitive function that may be early indicators of neurodegenerative disease but are challenging to measure with traditional data collection and analytic approaches. In another example using an app-based sensor, Zhang et al. (2025) recorded 30 seconds of ambient sound every few minutes for 5–6 days. These sound clips were used to passively collect linguistic features of speech in naturalistic environments to predict cognitive function. Finally, Al-Hammadi et al. (2025) also used sensor-based measures to capture real-world driving behavior. Using GPS location data from a datalogger installed in participants’ vehicles, they were able to improve ML-based predictions of preclinical AD by incorporating indices of socioeconomic deprivation from areas where individuals frequently drive. This demonstrates how digital technology and ML can be used to better understand the structural and social determinants of health. Overall, these studies emphasize the significant and innovative advancements in our understanding of aging and AD/ADRD that digital technology and AI are already yielding. The rapid growth of AI/ML over the past five years also ushered in a shortage of interdisciplinary expertise among data engineers/computer scientists knowledgeable about aging and disease, and gerontologists proficient in complex algorithms. This gap has led to methodological inconsistencies and incorrect assumptions between fields. In practice, approaches vary widely and lack consensus guidelines, resulting in issues such as model-data misalignment and unchecked biases (Hanna et al., 2025; Rudin, 2019). For example, algorithms trained on datasets in homogenous older adult samples tend to perform poorly on diverse populations, demonstrating suboptimal accuracy and fairness when applied to older cohorts (Das & Dhillon, 2023). Similarly, ML tools in aging research are often developed without rigorous validation against gerontological benchmarks, resulting in poor generalizability and reliability issues when tested in real-world settings. The absence of established training pipelines or best-practice standards has left AI applications for gerontology vulnerable to biases and methodological flaws that domain-specific experts alone might overlook (Navarro et al., 2021). Cross-disciplinary curricula and training programs that combine AI techniques with gerontological content can enhance fluency and bridge the gap between computational and aging research. Interdisciplinary research centers can institutionalize collaboration, bringing computer scientists, statisticians, clinicians, and aging experts together under a common umbrella. Equally important is developing consistent research standards tailored to this intersection. The research community needs to establish guidelines for data quality and fairness in aging-related AI (e.g., mandating inclusion of diverse age groups in training data) and promote benchmarks that align ML outcomes with clinically relevant aging outcomes (Thiyagalingam et al., 2022). Applying ML models to longitudinal data in neuroscience, psychology, and gerontology holds promise for the early detection of decline and the development of personalized interventions. Repeated imaging scans, electronic health records (EHRs), neuropsychological batteries, and digital phenotyping (e.g., smartphone or sensors) can be used to model aging trajectories. Frequent missing data across time and irregularly timed measurements are common in human studies, but can violate the assumptions of ML time-series methods, which are built on the assumption of complete samples (Carrasco-Ribelles et al., 2023; Lundberg et al., 2020). Imputation or specialized models for irregular time series are potential solutions, but they introduce complexity that may compromise replicability and can still lead to bias if missingness is non-random (Coupland et al., 2025). In longitudinal data, an individual’s repeated measures are statistically dependent, but many ML algorithms assume independent and identically distributed samples. Ignoring temporal dependencies can bias findings. Modeling strategies, such as mixed-effects models or sequence neural networks (recurrent or convolutional), are needed but require large data volumes. Aging trajectories are highly heterogeneous and non-linear, particularly in communities that are historically marginalized, minoritized, and underrepresented. This makes model generalizability a key challenge. Subtle fluctuations (e.g., in memory performance) might be meaningful in one person but noise in others. These complexities require careful temporal modeling and much larger datasets with diversity than are available in most aging studies. Sample representation and homogeneity are also concerns for ML models in gerontological research. Longitudinal cohorts often suffer from survivor bias (healthier individuals remain over time) and loss-to-follow-up bias (those with worsening conditions are more likely to drop out). Additionally, EHR-based studies may overlook individuals without regular access to healthcare. These sample biases mean an ML model may perform well on the study cohort but fail to generalize to other groups. ML models for longitudinal aging data risk overfitting due to high complexity and limited data. Some approaches (e.g., deep neural networks) are known for memorizing quirks of the training data that may not generalize to test and validation sets (Coupland et al., 2025). Generalizability requires testing on independent cohorts (ideally from different hospitals or demographics) to ensure the model does not overfit based on characteristics of a particular study. Moreover, data drift (changes in data over time and/or context) can degrade model performance. In EHR data, e.g., coding practices, devices, and diagnostic criteria evolve over time, which can lower model performance. Ensuring robust and generalizable models demands large, diverse datasets and ongoing validation. Finally, a fundamental limitation is that predictive ML on longitudinal data is usually correlational, not causal. Conventional ML models learn associations between variables in the data, without distinguishing between causal relationships and spurious correlations. Longitudinal data provide the temporal ordering required for causal insight; however, ML models themselves do not infer causality. While causal inference techniques exist (e.g., causal graphs, counterfactual frameworks), they are not automatically employed in typical ML pipelines. Applying ML/DL models to gerontological datasets requires caution in interpreting results, as a pattern does not necessarily equate to a putative causal factor. Achieving actionable insights for gerontological research, clinical practice, and precision medicine requires integrating domain knowledge and causal analysis with informed and intentional ML models. The anchoring editorial (Stoeckel et al., 2025) articulates a bold vision for AI-driven precision measurement in aging and AD/ADRD research. Their vision highlights the need for new tools that detect early cognitive, behavioral, and neuropsychiatric changes beyond traditional tests. Drawing on national initiatives (e.g., the PREPARE Challenge, Mobile Toolbox), they call for inclusive, multimodal, and ethically grounded innovations that integrate lived experience, representative data, and participatory design. This call serves as both a foundation and a catalyst, impelling researchers to rethink and transform measurement science through open, reproducible, and socially responsive AI practices. None. G.M.B. served as a coauthor in Al-Hammadi et al. (2025), included in the special issue, but was not involved in the review or decision for the article.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,014 | 0,062 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,003 | 0,002 |
| Études des sciences et des technologies | 0,002 | 0,011 |
| Communication savante | 0,008 | 0,015 |
| Science ouverte | 0,004 | 0,006 |
| Intégrité de la recherche | 0,004 | 0,007 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».