Identification of candidate domestication‐related genes with a systematic survey of loss‐of‐function mutations
Bibliographic record
Abstract
Domestication is an important key co-evolutionary process through which humans have extensively altered the genomic make-up and appearance of both plants and animals. The identification of domestication-related genes remains very arduous. In this study, we present a systematic analytical approach that harnesses two recent advances in genomics, whole-genome sequencing (WGS) and prediction of loss-of-function (LOF) mutations, to greatly facilitate the assembly of an enriched catalogue of domestication-related candidate genes. Using WGS data for 296 cultivated (Glycine max) and 64 wild soybean accessions, we identified 8699 LOF variants, and 116 genes that are uniquely fixed for one or more LOF allele(s) in domesticated soybeans. Existing soybean transcriptomic data led us to overcome analytical challenges associated with whole-genome duplications and to identify neo- or subfunctionalized genes. This systematic approach allowed us to identify 110 candidate domestication-related genes in an efficient and rapid way. This catalogue contains previously well characterized domestication genes in soybean, as well as some orthologs from other domesticated crop species. In addition, it comprises many promising candidate domestication genes. Overall, this collection of candidate domestication-related genes in soybean is almost twice as large as the sum of all previously reported candidate genes in all other crops. We believe this systematic approach could readily be used in wide range of species.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".