Assigning Phenologically Asynchronous Moths to Source Populations Using Individual Genotypes
Bibliographic record
Abstract
The spruce budworm (Choristoneura fumiferana; SBW) is a periodically outbreaking forest insect pest that affects the boreal forests of North America through extensive defoliation and tree mortality. Causes of widespread spatial synchrony of SBW outbreaks remain a key question in the ecology and management of this species. While the Moran effect (correlated favourable environmental conditions) and density-dependent dispersal (from epicentres of demographic explosions) have been proposed and supported as drivers of synchronised outbreaks, the relative contribution of long-distance dispersal is still poorly understood. In this study, we use a novel approach to distinguish resident from migrant moths and to assign migrants to likely source clusters with the goal of better characterising regional dispersal. First, we characterise the genetic diversity and structure of resident SBW larvae and three phenologically separated groups of moths over one flight season using Genotyping-by-Sequencing. Then, using a novel machine learning approach, we assign putative migrants to their likely source populations. We hypothesised that migrant moths and resident larvae would be genetically distinct and could be assigned to source populations. Our findings revealed complex patterns of moth dispersal and population differentiation within a single season, including two spatially overlapping genetic clusters. We observed subtle but significant genetic differences between resident larvae and migrant moths, supporting the hypothesis that long-distance dispersal contributes to outbreak dynamics and synchrony. These insights enhance our understanding of SBW population dynamics and suggest that effective management strategies, such as the Early Intervention Strategy (EIS), must account for the role of dispersal in mitigating the detrimental effects of major outbreaks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".