Predicting skeletal stature using ancient <scp>DNA</scp>
Bibliographic record
Abstract
Abstract Objectives Ancient DNA provides an opportunity to separate the genetic and environmental bases of complex traits by allowing direct estimation of genetic values in ancient individuals. Here, we test whether genetic scores for height in ancient individuals are predictive of their actual height, as inferred from skeletal remains. We estimate the contributions of genetic and environmental variables to observed phenotypic variation as a first step towards quantifying individual sources of morphological variation. Materials and methods We collected stature estimates and femur lengths from West Eurasian skeletal remains with published genome‐wide ancient DNA data ( n = 182, dating from 33,000–850 BP). We also recorded genetic sex, genetic ancestry, date and paleoclimate data for each individual, and δ 13 C and δ 15 N stable isotope values where available ( n = 69). We tested different methods of calculating polygenic scores, using summary statistics from four different genome wide association studies (GWAS) for height, and three methods for imputing missing genotypes. Results A polygenic score for height predicts 6.3% of the variance in femur length in our data ( n = 132, SD = 0.0069%, p = 0.001), controlling for sex, ancestry, and date. This is consistent with the predictive power of height PRS in present‐day populations and the low coverage of ancient samples. Comparatively, sex explains about 17% of the variance in femur length in our sample. Environmental effects also likely play a role in variation, independent of genetics, though with considerable uncertainty (longitude: R 2 = 0.033, SD = 0.008, p = 0.011). Genotype imputation did not improve polygenic prediction, and results varied based on the GWAS summary statistics we used. Discussion Polygenic scores explain a small but significant proportion of the variance in height in ancient individuals, though not enough to make useful predictions of individual phenotypes. However, environmental variables also contribute to phenotypic outcomes and understanding their interaction with direct genetic predictions will provide a framework with which to model how plasticity and genetic changes ultimately combine to drive adaptation and evolution.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".