MétaCan
Menu
Back to cohort
Record W2990822156 · doi:10.3389/fgene.2019.01192

LD-annot: A Bioinformatics Tool to Automatically Provide Candidate SNPs With Annotations for Genetically Linked Genes

2019· article· en· W2990822156 on OpenAlexafffund
Julien Prunier, Audrey Lemaçon, Alexandre Bastien, Mohsen Jafarikia, Ilga Porth, Claude Robert, Arnaud Droit

Bibliographic record

VenueFrontiers in Genetics · 2019
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Phylogenetic Studies
Canadian institutionsUniversity of GuelphCentre hospitalier universitaire de QuébecUniversité Laval
FundersNatural Resources CanadaGenome British ColumbiaCanadian Food Inspection AgencyGenome Canada
KeywordsSingle-nucleotide polymorphismGeneGeneticsBiologyComputational biologyComputer scienceCandidate geneBioinformaticsGenotype

Abstract

fetched live from OpenAlex

A multitude of model and non-model species studies have now taken full advantage of powerful high-throughput genotyping advances such as SNP arrays and genotyping-by-sequencing (GBS) technology to investigate the genetic basis of trait variation. However, due to incomplete genome coverage by these technologies, the identified SNPs are likely in linkage disequilibrium (LD) with the causal polymorphisms, rather than be causal themselves. In addition, researchers could benefit from annotations for the identified candidate SNPs and, simultaneously, for all neighboring genes in genetic linkage. In such case, LD extent estimation surrounding the candidate SNPs is required to determine the regions encompassing genes of interest. We describe here an automated pipeline, "LD-annot," designed to delineate specific regions of interest for a given experiment and candidate polymorphisms on the basis of LD extent, and furthermore, provide annotations for all genes within such regions. LD-annot uses standard file formats, bioinformatics tools, and languages to provide identifiers, coordinates, and annotations for genes in genetic linkage with each candidate polymorphism. Although the focus lies upon SNP arrays and GBS data as they are being routinely deployed, this pipeline can be applied to a variety of datasets as long as genotypic data are available for a high number of polymorphisms and formatted into a vcf file. A checkpoint procedure in the pipeline allows to test several threshold values for linkage without having to rerun the entire pipeline, thus saving the user computational time and resources. We applied this new pipeline to four different sample sets: two breeding populations GBS datasets, one within-pedigree SNP set coming from whole genome sequencing (WGS), and a very large multi-varieties SNP dataset obtained from WGS, representing variable sample sizes, and numbers of polymorphisms. LD-annot performed within minutes, even when very high numbers of polymorphisms are investigated and thus will efficiently assist research efforts aimed at identifying biologically meaningful genetic polymorphisms underlying phenotypic variation. LD-annot tool is available under a GPL license from https://github.com/ArnaudDroitLab/LD-annot.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.004
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Software · Consensus signal: none
Teacher disagreement score0.029
Threshold uncertainty score0.096

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0040.008
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0050.003
Science and technology studies0.0020.001
Scholarly communication0.0020.002
Open science0.0030.003
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0290.017

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.006
GPT teacher head0.224
Teacher spread0.218 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations16
Published2019
Admission routes2
Has abstractyes

Explore more

Same venueFrontiers in GeneticsSame topicGenomics and Phylogenetic StudiesFrench-language works237,207