Familial relatedness in genetic frontotemporal dementia cohorts: findings from the international Frontotemporal Dementia Prevention Initiative
Bibliographic record
Abstract
Abstract Background Given the rarity of genetic frontotemporal dementia (FTD), researchers across the world have come together to form the FTD Prevention Initiative (FPI) in an effort to improve prevention trials design. As this initiative begins to bring together large‐scale data from worldwide cohort series, including ALLFTD in North America and GENFI in Europe and Canada, it is critical for FPI to quantify the level of relatedness between all participants. Here we provide the most recent update to these ongoing analyses. Methods Genome‐wide SNP genotyping data from 1,684 ALLFTD and 568 GENFI participants was used to perform lineage analyses using PLINK. Briefly, QC was performed similarly in all datasets to remove individuals with low call rate and filter autosomal SNPs for missingness, frequency, and deviation from Hardy‐Weinberg equilibrium. Genetic ancestry was inferred by projecting genotyped samples into the principal components of the 1000 Genomes reference panel, using R package bigsnpr. Overlapping ALLFTD and GENFI genotyping data was then used, in a two‐stage approach, to calculate pairwise identity‐by‐descent (IBD) estimates and KING coefficients, followed by family‐network identification and pedigree reconstruction using PRIMUS. Results First, we calculated IBD estimates among all participants by restricting pairs to those with estimates>0.1875 (up to second‐degree relatives). Overall, we identified a total of 292 second‐degree family networks, including 168 ALLFTD and 120 GENFI families, mostly associated with pathogenic variants in the 3 major FTD‐causing genes. We also identified 4 family networks with participants enrolled in both the ALLFTD and GENFI series, as well as several multi‐site families within the ALLFTD consortium. This first, overall approach allowed us to predict close relationships even between individuals with different ancestral backgrounds, including at least 4 confirmed admixed families. More distant relationships were also detected within ALLFTD and GENFI by performing ancestry‐based analysis among participants with estimated European ancestry, using the KING‐robust algorithm. Conclusions These lineage analyses allowed us to identify, otherwise unknown, close (and distant) relatives from different study sites, as well as within the ALLFTD and GENFI series. This dataset will be a crucial resource to increase statistical accuracy and power in upcoming collaborative FPI studies.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.002 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".