Detection of Common Copy Number of Variants Underlying Selection Pressure in Middle Eastern Horse Breeds Using Whole-Genome Sequence Data
Bibliographic record
Abstract
Dareshouri, Arabian, and Akhal-Teke are 3 Middle Eastern horse breeds that have been selected for endurance and adaptation to harsh climates. Deciphering the genetic characteristics of these horses by tracing selection footprints and copy number of variations will be helpful in improving our understanding of equine breeds' development and adaptation. For this purpose, we sequenced the whole genome of 4 Dareshouri horses using Illumina Hiseq panels and compared them with publicly available whole-genome sequences of Arabian (n = 3) and Akhal-Teke (n = 3) horses. Three tests of FLK, hapFLK, and pooled heterozygosity were applied using a sliding window (window size = 100 kb, step size = 50 kb) approach to detect putative selection signals. Copy number variation analysis was applied to investigate copy number of variants (CNVs), and the results were used to suggest selection signatures involving CNVs. Whole-genome sequencing demonstrated 8 837 950 single-nucleotide polymorphisms (SNPs) in autosomal chromosomes. We suggested 58 genes and 3 quantitative trait loci, including some related to horse gait, insect bite hypersensitivity, and withers height, based on selective signals detected by adjusted P-value of Mahalanobis distance based on the rank-based P-values (Md-rank-P) method. We proposed 12 genomic regions under selection pressure involving CNVs that were previously reported to be associated with metabolism energy (SLC5A8), champagne dilution in horses (SLC36A1), and synthesis of polyunsaturated fatty acids (FAT2). Only 10 Middle Eastern horses were tested in this study; therefore, the conclusions are speculative. Our findings are useful to better understanding the evolution and adaptation of Middle Eastern horse breeds.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".