Patterns and processes in crop domestication: an historical review and quantitative analysis of 203 global food crops
Bibliographic record
Abstract
Summary Domesticated food crops are derived from a phylogenetically diverse assemblage of wild ancestors through artificial selection for different traits. Our understanding of domestication, however, is based upon a subset of well‐studied ‘model’ crops, many of them from the Poaceae family. Here, we investigate domestication traits and theories using a broader range of crops. We reviewed domestication information (e.g. center of domestication, plant traits, wild ancestors, domestication dates, domestication traits, early and current uses) for 203 major and minor food crops. Compiled data were used to test classic and contemporary theories in crop domestication. Many typical features of domestication associated with model crops, including changes in ploidy level, loss of shattering, multiple origins, and domestication outside the native range, are less common within this broader dataset. In addition, there are strong spatial and temporal trends in our dataset. The overall time required to domesticate a species has decreased since the earliest domestication events. The frequencies of some domestication syndrome traits (e.g. nonshattering) have decreased over time, while others (e.g. changes to secondary metabolites) have increased. We discuss the influences of the ecological, evolutionary, cultural and technological factors that make domestication a dynamic and ongoing process. Contents Summary 29 I. Introduction 30 II. Key concepts and definitions 30 III. Methods of review and analysis 35 IV. Trends identified from the review of 203 crops 37 V. Life cycle 38 VI. Ploidy level 40 VII. Reproductive strategies 42 VIII. The domestication syndrome 42 IX. Spatial and temporal trends 42 X. Utilization of plant parts 44 XI. Conclusions 44 Acknowledgements 45 References 45
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.013 | 0.021 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".