Understanding the evolutionary origin and ancestral composition of honey bee (Apis mellifera) populations.
Bibliographic record
Abstract
The honey bee, Apis mellifera, is arguably the most important managed pollinator globally. Yet despite its economic and ecological importance, there are still several unknowns regarding the species ancestral origin and ancestral complexity. Understanding the genetic composition of native and managed honey bee colonies is imperative for resolving the species life history and elucidating how ancestry may inform management strategies. In this dissertation, I take a deep dive into the evolutionary origins of Apis mellifera and learn how ancestral complexity has shaped the composition of contemporary populations. In Chapter two, I settle a long-standing debate about the ancestral origins of the species. I find that Apis mellifea diverged out of Western Asia via at least three colonization routes, which resulted in the evolution of at least seven genetically distinct lineages. Interesting, I find that these lineages were able to adapt to their current distribution by repeated selection among a core set of genes. In Chapter three, I take a closer look at the genetic complexity of managed Canadian honey bees by estimating the ancestral composition of colonies using the genomic dataset from Chapter two. I find that patterns of ancestry differ between Canadian provinces, and that admixture correlates strongly with levels of genetic diversity. Interestingly, I find that genomic intervals with elevated levels of admixture segregate non-randomly in the genome and are associated with genes related to parasite and xenobiotic tolerance. Though admixture may bear advantages for managed colonies, admixture among honey bee is not always valued. In Chapter four and five I make use of the ancestral composition of invasive Africanized honey bees to develop assays to identify and track populations. This was achieved using machine learning models to choose the most informative single nucleotide polymorphisms (Chapter 4) and insertion-deletion (Chapter 5) markers that best delineate Africanized genetics from managed European colonies. My research addresses many gaps in our understanding of honey bee origins and ancestral complexity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".