The effect of using cow genomic information on accuracy and bias of genomic breeding values in a simulated Holstein dairy cattle population
Bibliographic record
Abstract
Using cow data in the training population is attractive as a way to mitigate bias due to highly selected training bulls and to implement genomic selection for countries with no or limited proven bull data. However, one potential issue with cow data is a bias due to the preferential treatment. The objectives of this study were to (1) investigate the effect of including cow genotype and phenotype data into the training population on accuracy and bias of genomic predictions and (2) assess the effect of preferential treatment for different proportions of elite cows. First, a 4-pathway Holstein dairy cattle population was simulated for 2 traits with low (0.05) and moderate (0.3) heritability. Then different numbers of cows (0, 2,500, 5,000, 10,000, 15,000, or 20,000) were randomly selected and added to the training group composed of different numbers of top bulls (0, 2,500, 5,000, 10,000, or 15,000). Reliability levels of de-regressed estimated breeding values for training cows and bulls were 30 and 75% for traits with low heritability and were 60 and 90% for traits with moderate heritability, respectively. Preferential treatment was simulated by introducing upward bias equal to 35% of phenotypic variance to 5, 10, and 20% of elite bull dams in each scenario. Two different validation data sets were considered: (1) all animals in the last generation of both elite and commercial tiers (n = 42,000) and (2) only animals in the last generation of the elite tier (n = 12,000). Adding cow data into the training population led to an increase in accuracy (r) and decrease in bias of genomic predictions in all considered scenarios without preferential treatment. The gain in r was higher for the low heritable trait (from 0.004 to 0.166 r points) compared with the moderate heritable trait (from 0.004 to 0.116 r points). The gain in accuracy in scenarios with a lower number of training bulls was relatively higher (from 0.093 to 0.166 r points) than with a higher number of training bulls (from 0.004 to 0.09 r points). In this study, as expected, the bull-only reference population resulted in higher accuracy compared with the cow-only reference population of the same size. However, the cow reference population might be an option for countries with a small-scale progeny testing scheme or for minor breeds in large counties, and for traits measured only on a small fraction of the population. The inclusion of preferential treatment to 5 to 20% of the elite cows led to an adverse effect on both accuracy and bias of predictions. When preferential treatment was present, random selection of cows did not reduce the effect of preferential treatment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.021 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".