Sequencing of <i>Camelina neglecta</i>, a diploid progenitor of the hexaploid oilseed <i>Camelina sativa</i>
Bibliographic record
Abstract
Camelina neglecta is a diploid species from the genus Camelina, which includes the versatile oilseed Camelina sativa. These species are closely related to Arabidopsis thaliana and the economically important Brassica crop species, making this genus a useful platform to dissect traits of agronomic importance while providing a tool to study the evolution of polyploids. A highly contiguous chromosome-level genome sequence of C. neglecta with an N50 size of 29.1 Mb was generated utilizing Pacific Biosciences (PacBio, Menlo Park, CA) long-read sequencing followed by chromosome conformation phasing. Comparison of the genome with that of C. sativa shows remarkable coincidence with subgenome 1 of the hexaploid, with only one major chromosomal rearrangement separating the two. Synonymous substitution rate analysis of the predicted 34 061 genes suggested subgenome 1 of C. sativa directly descended from C. neglecta around 1.2 mya. Higher functional divergence of genes in the hexaploid as evidenced by the greater number of unique orthogroups, and differential composition of resistant gene analogs, might suggest an immediate adaptation strategy after genome merger. The absence of genome bias in gene fractionation among the subgenomes of C. sativa in comparison with C. neglecta, and the complete lack of fractionation of meiosis-specific genes attests to the neopolyploid status of C. sativa. The assembled genome will provide a tool to further study genome evolution processes in the Camelina genus and potentially allow for the identification and exploitation of novel variation for Camelina crop improvement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".