From algorithms to zero-shot learning: how artificial intelligence is redefining gerontological research
Bibliographic record
Abstract
By 2050, the global population of adults (≥60 years) is expected to double, to 2.1 billion (World Health Organization, 2024). To support older adults’ independence and participation in society, communities will need to create age-friendly and inclusive healthcare systems that manage multiple chronic conditions, expand long-term and dementia care, and redesign housing and transportation to ensure accessibility (Chum et al., 2022; Rowe et al., 2016). The intersection of artificial intelligence (AI) and gerontological science is accelerating our understanding of common conditions and management of aging-related cognitive and behavioral changes in the aging population (Chen et al., 2023). AI has emerged as a transformative tool in gerontological research, enabling the analysis of large-scale health data to enhance diagnostics, predictive risk modeling, and overall care delivery for older adults (Bernal et al., 2024). AI algorithms hold promise in supporting early disease detection and diagnosis, identifying conditions like delirium, and predicting adverse events like falls through machine learning (O’Connor et al., 2022; Pagali et al., 2023). Additionally, AI enables continuous monitoring via wearable sensors and smart home systems, allowing for tailored interventions and improved resource allocation in healthcare. Researchers are also leveraging AI to analyze longitudinal aging datasets, such as medical imaging and electronic health records, to map trajectories of cognitive decline and other age-related changes, offering new insights into dementia progression and functional aging patterns (Kourtis et al., 2019). This special issue of the Psychological Sciences section of The Journals of Gerontology, Series B, explores how metrics, including digital biomarkers and phenotypes, can improve early detection, disease monitoring, and intervention in aging populations. It highlights the innovative use of AI techniques ranging from traditional machine learning (ML) and deep learning (DL) to natural language processing and computer vision in advancing the precision measurement of cognitive, psychological, and social functioning among older adults. Studies included in this special issue fall into two broad categories. The first category comprises investigations focused on the use of AI/ML to analyze or improve the classification and prediction of aging and disease-related characteristics based on traditional, yet often high-dimensional, data types (e.g., demographic, clinical, and neuropsychological data). These analyses demonstrate how AI/ML can be leveraged to extract new insights from existing data, where traditional approaches have been less well-suited for data-driven analyses. For example, Jiang & Yu (2025) used ML to identify factors related to social resilience in older adults across three datasets, thereby supporting the prioritization of interventions to improve the social adaptability of older adults. Other papers focus on using AI/ML to improve the prediction of health and Alzheimer’s disease and related dementia (AD/ADRD)-related characteristics and outcomes based on existing data. Petersen et al. (2025) developed risk scores for predicting amyloid and tau biomarkers using ML (AutoScore) applied to data collected as part of the Alzheimer’s Disease Neuroimaging dataset. This approach enabled hierarchical combinations of demographic, neuropsychological, genetic, and imaging characteristics for prediction, in a manner that could enhance participant screening for clinical trials. Other contributions leveraged computer vision algorithms to better capture features of traditional neuropsychological task data. Hu et al. (2025) used deep learning neural network models to improve dementia classification from clock drawing images collected as part of the National Health and Aging Trends Study. In addition to improving efficiency, this approach allowed the development of continuous scores that were better able to balance sensitivity and specificity. Other studies in this special issue also highlight how AI/ML can be applied to prediction based on time-series data. Using transformer-based models, Weiss et al. (2025) achieved a twofold improvement in (next-wave) mortality prediction using data from the Health and Retirement Study collected over 29 years. In addition to advances in prediction, this special issue includes work validating approaches to ML-based classification. Nguyen & Chang (2025) compared results from different data-driven clustering algorithms (nonnegative matrix factorization and model-based clustering) with distinct heuristics, applied to demographic, clinical, neuropsychological, and brain morphometric data. Their analysis supports the comparability of these clustering algorithms for generating convergent insights. The second category of contribution included in this special issue involves the application of AI/ML to data generated from current and next-generation digital technologies, including passive data collected from sensors and smartphone-based digital cognitive assessments. These data are often temporally dense and may not be amenable to many traditional data analytic approaches; however, they can be used to generate novel digital phenotypes that capture essential information about health and behavior. Oravecz et al. (2025) utilized high-frequency longitudinal data from smartphone-based cognitive assessments to identify computational cognitive markers of long-term forgetting, learning, and within-person variability related to cognitive aging and mild cognitive impairment. Importantly, these markers reflect changes in cognitive function that may be early indicators of neurodegenerative disease but are challenging to measure with traditional data collection and analytic approaches. In another example using an app-based sensor, Zhang et al. (2025) recorded 30 seconds of ambient sound every few minutes for 5–6 days. These sound clips were used to passively collect linguistic features of speech in naturalistic environments to predict cognitive function. Finally, Al-Hammadi et al. (2025) also used sensor-based measures to capture real-world driving behavior. Using GPS location data from a datalogger installed in participants’ vehicles, they were able to improve ML-based predictions of preclinical AD by incorporating indices of socioeconomic deprivation from areas where individuals frequently drive. This demonstrates how digital technology and ML can be used to better understand the structural and social determinants of health. Overall, these studies emphasize the significant and innovative advancements in our understanding of aging and AD/ADRD that digital technology and AI are already yielding. The rapid growth of AI/ML over the past five years also ushered in a shortage of interdisciplinary expertise among data engineers/computer scientists knowledgeable about aging and disease, and gerontologists proficient in complex algorithms. This gap has led to methodological inconsistencies and incorrect assumptions between fields. In practice, approaches vary widely and lack consensus guidelines, resulting in issues such as model-data misalignment and unchecked biases (Hanna et al., 2025; Rudin, 2019). For example, algorithms trained on datasets in homogenous older adult samples tend to perform poorly on diverse populations, demonstrating suboptimal accuracy and fairness when applied to older cohorts (Das & Dhillon, 2023). Similarly, ML tools in aging research are often developed without rigorous validation against gerontological benchmarks, resulting in poor generalizability and reliability issues when tested in real-world settings. The absence of established training pipelines or best-practice standards has left AI applications for gerontology vulnerable to biases and methodological flaws that domain-specific experts alone might overlook (Navarro et al., 2021). Cross-disciplinary curricula and training programs that combine AI techniques with gerontological content can enhance fluency and bridge the gap between computational and aging research. Interdisciplinary research centers can institutionalize collaboration, bringing computer scientists, statisticians, clinicians, and aging experts together under a common umbrella. Equally important is developing consistent research standards tailored to this intersection. The research community needs to establish guidelines for data quality and fairness in aging-related AI (e.g., mandating inclusion of diverse age groups in training data) and promote benchmarks that align ML outcomes with clinically relevant aging outcomes (Thiyagalingam et al., 2022). Applying ML models to longitudinal data in neuroscience, psychology, and gerontology holds promise for the early detection of decline and the development of personalized interventions. Repeated imaging scans, electronic health records (EHRs), neuropsychological batteries, and digital phenotyping (e.g., smartphone or sensors) can be used to model aging trajectories. Frequent missing data across time and irregularly timed measurements are common in human studies, but can violate the assumptions of ML time-series methods, which are built on the assumption of complete samples (Carrasco-Ribelles et al., 2023; Lundberg et al., 2020). Imputation or specialized models for irregular time series are potential solutions, but they introduce complexity that may compromise replicability and can still lead to bias if missingness is non-random (Coupland et al., 2025). In longitudinal data, an individual’s repeated measures are statistically dependent, but many ML algorithms assume independent and identically distributed samples. Ignoring temporal dependencies can bias findings. Modeling strategies, such as mixed-effects models or sequence neural networks (recurrent or convolutional), are needed but require large data volumes. Aging trajectories are highly heterogeneous and non-linear, particularly in communities that are historically marginalized, minoritized, and underrepresented. This makes model generalizability a key challenge. Subtle fluctuations (e.g., in memory performance) might be meaningful in one person but noise in others. These complexities require careful temporal modeling and much larger datasets with diversity than are available in most aging studies. Sample representation and homogeneity are also concerns for ML models in gerontological research. Longitudinal cohorts often suffer from survivor bias (healthier individuals remain over time) and loss-to-follow-up bias (those with worsening conditions are more likely to drop out). Additionally, EHR-based studies may overlook individuals without regular access to healthcare. These sample biases mean an ML model may perform well on the study cohort but fail to generalize to other groups. ML models for longitudinal aging data risk overfitting due to high complexity and limited data. Some approaches (e.g., deep neural networks) are known for memorizing quirks of the training data that may not generalize to test and validation sets (Coupland et al., 2025). Generalizability requires testing on independent cohorts (ideally from different hospitals or demographics) to ensure the model does not overfit based on characteristics of a particular study. Moreover, data drift (changes in data over time and/or context) can degrade model performance. In EHR data, e.g., coding practices, devices, and diagnostic criteria evolve over time, which can lower model performance. Ensuring robust and generalizable models demands large, diverse datasets and ongoing validation. Finally, a fundamental limitation is that predictive ML on longitudinal data is usually correlational, not causal. Conventional ML models learn associations between variables in the data, without distinguishing between causal relationships and spurious correlations. Longitudinal data provide the temporal ordering required for causal insight; however, ML models themselves do not infer causality. While causal inference techniques exist (e.g., causal graphs, counterfactual frameworks), they are not automatically employed in typical ML pipelines. Applying ML/DL models to gerontological datasets requires caution in interpreting results, as a pattern does not necessarily equate to a putative causal factor. Achieving actionable insights for gerontological research, clinical practice, and precision medicine requires integrating domain knowledge and causal analysis with informed and intentional ML models. The anchoring editorial (Stoeckel et al., 2025) articulates a bold vision for AI-driven precision measurement in aging and AD/ADRD research. Their vision highlights the need for new tools that detect early cognitive, behavioral, and neuropsychiatric changes beyond traditional tests. Drawing on national initiatives (e.g., the PREPARE Challenge, Mobile Toolbox), they call for inclusive, multimodal, and ethically grounded innovations that integrate lived experience, representative data, and participatory design. This call serves as both a foundation and a catalyst, impelling researchers to rethink and transform measurement science through open, reproducible, and socially responsive AI practices. None. G.M.B. served as a coauthor in Al-Hammadi et al. (2025), included in the special issue, but was not involved in the review or decision for the article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.062 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.002 | 0.011 |
| Scholarly communication | 0.008 | 0.015 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.004 | 0.007 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".