Challenges in predicting brain age with high-dimensional neuroimaging data: insights from quantile models and morphometricity estimation
Bibliographic record
Abstract
Abnormal brain evolution is linked to various neurological conditions, including Autism Spectrum Disorder and Alzheimer's Disease.This anomaly can be quantified using "brain age," which identifies the age group within a healthy cohort whose brain resembles that of the individual.A significant difference between predicted brain age and chronological age suggests delayed brain development or accelerated aging.Current brain age prediction models mainly focus on the average brain age within a population, based on imaging profiles at a specific chronological age.However, little attention is given to brain age distributions and the behavior of various quantiles.Inspired by classic infant growth charts, my primary goal is to develop a comprehensive growth model for the human brain that provides predictions across multiple quantiles.Such a model would enhance our ability to quantify deviations in an individual's brain development from the normal distribution.Given the complexity of structural brain imaging data, flexible machine learning models are necessary.My research highlights a lack of suitable methods for validating and evaluating these complex models, addressing two key shortcomings in my thesis.First, I explore the proportion of phenotypic variance explained by features, often used as a benchmark for the maximum achievable prediction accuracy of statistical models.This proportion, known as "morphometricity" in brain morphology, is typically estimated using linear mixed-effects models.Through extensive simulations, I show that choices of hyperparameters(e.g.kernel and bandwidth) significantly impact morphometricity estimates.Conventional likelihood-based model selection methods tend to favor i Dr. Celia Greenwood and Dr. Jean-Baptiste Poline.Your invaluable insights, unwavering patience, and encouragement have been instrumental in shaping my research and guiding me through the challenges of this project.I am grateful for the freedom you allowed me to explore my own ideas, and for your unwavering belief in my work.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.051 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".