MétaCan
Menu
Retour à la cohorte
Enregistrement W7046778010

Developing a reproducible bioinformatics workflow for canine inherited retinal disease

2023· other· en· W7046778010 sur OpenAlexaboutno aff

Notice bibliographique

RevueKTH Publication Database DiVA (KTH Royal Institute of Technology) · 2023
Typeother
Langueen
DomainePhysics and Astronomy
ThématiqueMagnetic confinement fusion research
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésWorkflowDiseaseMendelian inheritancePipeline (software)ZebrafishModel organismGeneGenomeRetinal
DOInon disponible

Résumé

récupéré en direct d'OpenAlex

Inherited Retinal Degenerations (IRDs) are a heterogenous group of diseases which lead to vision impairment and can be found both in humans and in dogs. About 1 in 1,380 humans is estimated to suffer from an autosomal recessive IRD, which would be 5.5 million people worldwide, and many more are estimated to be unaffected carriers. This makes autosomal recessive IRDs likely the most common group of Mendelian diseases in humans. Today, about 300 genetic mutations have been connected to cause retinal diseases in humans. Whilst in dogs only 32 genes have been identified, numerous eye conditions have been described where the genetic cause has not yet been identified. This suggests that there are much more genetic causes to discover in the dog genome. Additionally, the dog serves well as a model organism to investigate IRDs as it is sharing morphological and genetic similarities with humans. For these reasons, proper software, a canine reference genome of high quality, and smart implementation of bioinformatic tools and methods are a big advantage to increase chances of finding new causative genetic variants and subsequently enable faster detection of possible preventions of the disease or at least alleviating its symptoms via early diagnosis. In this project, a pre-existing pipeline consisting of Bash scripts was stepwise improved with the goal to increase its efficiency. After controlling whether previous data could still be reproduced with the old pipeline in a first step, the software was exchanged to more updated versions in a second step. A main change was the replacement of the mapping tool Burrows-Wheeler Aligner (BWA) from bwa mem to bwa-mem2 mem, and the update of deprecated Genome Analysis Toolkit (GATK) 3.7 to version 4.3 or 4.4. Thirdly, the scripts were adapted from using the older canine reference genome CanFam3.1 to CanFam4. In a fourth step, for automatization and fastening the running time, the pipeline steps were implemented into the workflow management system Nextflow. Additionally, this step was partly aiming to make the pipeline in concordance with the FAIR-principles. All steps were tested on the same test data set, a Labrador retriever family trio, in which one genetic cause for a canine form of the IRD Stargardt disease in a previous study had been detected, namely an insertion in the ABCA4 gene. Lastly, the workflow was also tested on a second data set of a novel IRD of unknown genetic origin on two sibling pairs of Chinese Crested Dogs (CCR). The adjustment of the pipeline shows similar results regarding the change of mapping tool. Introducing the new reference genome revealed a drop of average coverage by one read average for when using CanFam4, while other results were similar. Using the new reference genome increased the number of unknown variants compared to findings with CanFam3.1. However, the known causative variant for the canine form of Stargardt disease, an insertion in ABCA4 gene, could be found in all cases. The run with Nextflow produced identical results to when the respective steps were run with Bash scripts, but it reduced the running time. Running the workflow on the new data set (CCR) and subsequent annotation and filtering indicate new candidates which could be further investigated as a potential cause for this currently unknown cause for an IRD.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,010
score de la tête « metaresearch » (Gemma)0,016
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: Reproductibilité · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Méthodes · Signal consensuel: Méthodes
Score de désaccord entre enseignants0,990
Score d'incertitude au seuil0,053

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0100,016
Méta-épidémiologie (sens strict)0,0020,001
Méta-épidémiologie (sens large)0,0020,003
Bibliométrie0,0030,002
Études des sciences et des technologies0,0020,001
Communication savante0,0050,002
Science ouverte0,0030,004
Intégrité de la recherche0,0010,003
Charge utile insuffisante (le modèle a refusé de juger)0,0130,020

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,034
Tête enseignante GPT0,302
Écart entre enseignants0,267 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Devis d'étudeSans objet
DomaineReproductibilité
GenreMéthodes

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2023
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueKTH Publication Database DiVA (KTH Royal Institute of Technology)Même sujetMagnetic confinement fusion researchTravaux en français237 207