A data-driven disease progression model of fluid biomarkers in genetic frontotemporal dementia
Bibliographic record
Abstract
Several CSF and blood biomarkers for genetic frontotemporal dementia have been proposed, including those reflecting neuroaxonal loss (neurofilament light chain and phosphorylated neurofilament heavy chain), synapse dysfunction [neuronal pentraxin 2 (NPTX2)], astrogliosis (glial fibrillary acidic protein) and complement activation (C1q, C3b). Determining the sequence in which biomarkers become abnormal over the course of disease could facilitate disease staging and help identify mutation carriers with prodromal or early-stage frontotemporal dementia, which is especially important as pharmaceutical trials emerge. We aimed to model the sequence of biomarker abnormalities in presymptomatic and symptomatic genetic frontotemporal dementia using cross-sectional data from the Genetic Frontotemporal dementia Initiative (GENFI), a longitudinal cohort study. Two-hundred and seventy-five presymptomatic and 127 symptomatic carriers of mutations in GRN, C9orf72 or MAPT, as well as 247 non-carriers, were selected from the GENFI cohort based on availability of one or more of the aforementioned biomarkers. Nine presymptomatic carriers developed symptoms within 18 months of sample collection ('converters'). Sequences of biomarker abnormalities were modelled for the entire group using discriminative event-based modelling (DEBM) and for each genetic subgroup using co-initialized DEBM. These models estimate probabilistic biomarker abnormalities in a data-driven way and do not rely on previous diagnostic information or biomarker cut-off points. Using cross-validation, subjects were subsequently assigned a disease stage based on their position along the disease progression timeline. CSF NPTX2 was the first biomarker to become abnormal, followed by blood and CSF neurofilament light chain, blood phosphorylated neurofilament heavy chain, blood glial fibrillary acidic protein and finally CSF C3b and C1q. Biomarker orderings did not differ significantly between genetic subgroups, but more uncertainty was noted in the C9orf72 and MAPT groups than for GRN. Estimated disease stages could distinguish symptomatic from presymptomatic carriers and non-carriers with areas under the curve of 0.84 (95% confidence interval 0.80-0.89) and 0.90 (0.86-0.94) respectively. The areas under the curve to distinguish converters from non-converting presymptomatic carriers was 0.85 (0.75-0.95). Our data-driven model of genetic frontotemporal dementia revealed that NPTX2 and neurofilament light chain are the earliest to change among the selected biomarkers. Further research should investigate their utility as candidate selection tools for pharmaceutical trials. The model's ability to accurately estimate individual disease stages could improve patient stratification and track the efficacy of therapeutic interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".