SNP-based analysis reveals unexpected features of genetic diversity, parental contributions and pollen contamination in a white spruce breeding program
Bibliographic record
Abstract
Accurate monitoring of genetic diversity levels of seedlots and mating patterns of parents from seed orchards are crucial to ensure that tree breeding programs are long-lasting and will deliver anticipated genetic gains. We used SNP genotyping to characterize founder trees, five bulk seed orchard seedlots, and trees from progeny trials to assess pollen contamination and the impact of severe roguing on genetic diversity and parental contributions in a first-generation open-pollinated white spruce clonal seed orchard. After severe roguing (eliminating 65% of the seed orchard trees), we found a slight reduction in the Shannon Index and a slightly negative inbreeding coefficient, but a sharp decrease in effective population size (eightfold) concomitant with sharp increase in coancestry (eightfold). Pedigree reconstruction showed unequal parental contributions across years with pollen contamination levels between 12 and 51% (average 27%) among seedlots, and 7-68% (average 30%) among individual genotypes within a seedlot. These contamination levels were not correlated with estimates obtained using pollen flight traps. Levels of pollen contamination also showed a Pearson's correlation of 0.92 with wind direction, likely from a pollen source 1 km away from the orchard under study. The achievement of 5% genetic gain in height at rotation through eliminating two-thirds of the orchard thus generated a loss in genetic diversity as determined by the reduction in effective population size. The use of genomic profiles revealed the considerable impact of roguing on genetic diversity, and pedigree reconstruction of full-sib families showed the unanticipated impact of pollen contamination from a previously unconsidered source.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".