A weakly structured stem for human origins in Africa
Bibliographic record
Abstract
Abstract While it is now broadly accepted that Homo sapiens originated within Africa, considerable uncertainty surrounds specific models of divergence and migration across the continent. Progress is hampered by a paucity of fossil and genomic data, as well as variability in prior divergence time estimates. Here we use linkage disequilibrium and diversity-based statistics, optimized for rapid, complex demographic inference to discriminate among such models. We infer detailed demographic models for populations across Africa, including representatives from eastern and western groups, as well as 44 newly whole-genome sequenced individuals from the Nama (Khoe-San). Despite the complexity of African population history, contemporary population structure dates back to Marine Isotope Stage (MIS) 5. The earliest population divergence among contemporary populations occurs 120-135ka, between the Khoe-San and other groups. Prior to the divergence of contemporary African groups, we infer long-lasting structure between two or more weakly differentiated ancestral Homo populations connected by gene flow over hundreds of thousands of years (i.e. a weakly structured stem). We find that weakly structured stem models provide more likely explanations of polymorphism that had previously been attributed to contributions from archaic hominins in Africa. In contrast to models with archaic introgression, we predict that fossil remains from coexisting ancestral populations should be morphologically similar. Despite genetic similarity between these populations, an inferred 1–4% of genetic differentiation among contemporary human populations can be attributed to genetic drift between stem populations. We show that model misspecification explains variation in previous divergence time estimates and argue that studying a suite of models is key to robust inferences about deep history.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".