Integrating <scp>QTL</scp> and expression <scp>QTL</scp> of <scp>PigGTEx</scp> to improve the accuracy of genomic prediction for small population in Yorkshire pigs
Bibliographic record
Abstract
The size of the reference population and sufficient phenotypic records are crucial for the accuracy of genomic selection. However, for small-to-medium-sized pig farms or breeds with limited population sizes, conducting genomic breeding programs presents significant challenges. In this study, 2295 Yorkshire pigs were selected from three distinct regions, including 1500 from an American line, 500 from a Canadian line, and 295 from a Danish line. All populations were genotyped using the GeneSeek 50K GGP Porcine HD chip. To enhance genomic selection accuracy, we proposed strategies that combined multiple populations and leveraged multi-omics prior information. Cis-QTL from the PigGTEx database and QTL identified through genome-wide association studies were incorporated into the genomic feature best linear unbiased prediction (GFBLUP) model to predict the ADG100 and the BF100 traits. Results demonstrated that combining multiple populations effectively improved prediction accuracy for small population, accuracy for ADG100 increased by an average of 0.29 and accuracy for BF100 by 0.05. The GFBLUP model, which integrates biological priors, showed some improvements in prediction accuracy for the BF100 trait. Specifically, for the small population, accuracy increased by 0.09 in Scheme 1, where each population size was predicted independently. In Scheme 3, where the large population was used as a reference group to predict the small population, accuracy increased by 0.03. However, the GFBLUP model did not provide additional benefits in predicting the ADG100 trait. These findings offer effective strategies for genetic improvement in developing regions and highlight the potential of multi-omics integration to enhance prediction models.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".