Implementation of Bayesian methods to identify SNP and haplotype regions with transmission ratio distortion across the whole genome: TRDscan v.1.0
Bibliographic record
Abstract
Realized deviations from the expected Mendelian inheritance of alleles from heterozygous parents have been previously reported in a broad range of organisms (i.e., transmission ratio distortion; TRD). Various biological mechanisms affecting gametes, embryos, fetuses, or even postnatal offspring can produce patterns of TRD. However, knowledge about its prevalence and potential causes in livestock species is still scarce. Specific Bayesian models have been recently developed for the analyses of TRD for biallelic loci, which accommodated a wide range of population structures, enabling TRD investigation in livestock populations. The parameterization of these models is flexible and allows the study of overall (parent-unspecific) TRD and sire- and dam-specific TRD. This research aimed at deriving Bayesian models for fitting TRD on the basis of haplotypes, testing the models for both haplotype- and SNP-based methods in simulated data and actual Holstein genotypes, and developing a specific software for TRD analyses. Results obtained on simulated data sets showed that the statistical power of the analysis increased with sample size of trios (n), proportion of heterozygous parents, and the magnitude of the TRD. On the other hand, the statistical power to detect TRD decreased with the number of alleles at each loci. Bayesian analyses showed a strong Pearson correlation coefficient (≥0.97) between simulated and estimated TRD that reached the significance level of Bayes factor ≥10 for both single-marker and haplotype analyses when n ≥ 25. Moreover, the accuracy in terms of the mean absolute error decreased with the increase of the sample size and increased with the number of alleles at each loci. Using real data (55,732 genotypes of Holstein trios), SNP- and haplotype-based distortions were detected with overall TRD, sire-TRD, or dam-TRD, showing different magnitudes of TRD and statistical relevance. Additionally, the haplotype-based method showed more ability to capture TRD compared with individual SNP. To discard possible random TRD in real data, an approximate empirical null distribution of TRD was developed. The program TRDscan v.1.0 was written in Fortran 2008 language and provides a powerful statistical tool to scan for TRD regions across the whole genome. This developed program is freely available at http://www.casellas.info/files/TRDscan.zip.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.010 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.004 | 0.002 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.024 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".