MétaCan
Menu
← Retour à la cohorte
Enregistrement W3213958976 · doi:10.1182/blood-2021-147548

An Open-Source Toolkit That Powers the Genome-Wide Analysis of Mature B-Cell Lymphomas (GAMBL) Project

2021· article· en· W3213958976 sur OpenAlexaff
Kostiantyn Dreval, Bruno M. Grande, Helena Winata, Jasper Wong, Lakshay Sethi, Christopher Rushton, Prasath Pararajalingam, Sarah E. Arthur, Lauren C. Chong, Brett Collinge, Krysta M. Coyle, Manuela Cruz, Stacy Hung, Shaghayegh Soudi, Nicole Thomas, Christian Steidl, David W. Scott, Ryan D. Morin, Laura K. Hilton

Notice bibliographique

RevueBlood · 2021
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueCancer Genomics and Diagnostics
Établissements canadiensCanada's Michael Smith Genome Sciences CentreSpinal Cord Injury BCSimon Fraser University
Organismes subventionnairesnon disponible
Mots-clésExomeGenomeComputational biologyBiologyLymphomaGenomicsExome sequencingBioinformaticsGeneGeneticsMutation

Résumé

récupéré en direct d'OpenAlex

Abstract Introduction: Genome- and transcriptome-wide analyses continue to enhance our understanding of the molecular pathogenesis of cancer. In lymphomas, this has enabled the identification of hundreds of recurrently mutated genes, highlighting genetic heterogeneity and relationships both within and among clinical entities. While the growing availability of lymphoma genomic data sets can be leveraged to integrate genomic analyses into diagnostic testing and clinical trials, the ability to rapidly process genomic data sets in a reproducible manner serves as a barrier to this goal. To this end, we developed a suite of tools Lymphoid Cancer Research modules (LCR-modules) to facilitate the discovery of novel drivers and molecular features in lymphoma cancers and perform quantitative comparisons between disease entities. We demonstrate here how this toolkit enabled a meta-analysis of lymphoma genomic data involving genome-wide profiles of 3330 patients. Methods: We assembled a collection of whole genome, whole exome, and RNA sequencing data from a combination of controlled-access repositories and ongoing projects at BC Cancer. The scope of genomic analysis of mature B-cell lymphomas (GAMBL) project includes cell lines and patient tumors from all common mature B cell neoplasms, comprising a total of 4612 samples from 3330 patients. To facilitate the project, we developed a suite of open-source and custom bioinformatics tools (https://github.com/LCR-BCCRC/lcr-modules) that leverages the Snakemake workflow management system and includes lymphoma-centric modules for the discovery and annotation of common mutation types, analysis of B-cell receptor repertoires and discovery of novel aSHM targets and relevant non-coding mutations, and RNA-seq analysis with batch correction and normalization. Individual modules are configured to create an automated, scalable, and reproducible workflow that runs each step as dictated by the availability of new data. The cohort-level integrative analysis and comparisons across entities are handled by our custom R package GAMBLR, which facilitates open-ended data analysis and custom visualizations. Results: Simple somatic mutations (SSM) were detected using a workflow that utilizes four algorithms to identify high-confidence variants with validated default thresholds for filtering of germline variants and common FFPE-associated artifacts, allowing for processing of samples without matched normal tissue. This automated and reproducible workflow facilitated the discovery of novel genes significantly mutated across lymphomas and broadened our understanding of the scope of aberrant somatic hypermutation (aSHM) and other non-coding mutations. Specifically, HNRNPU, STAT3, TFAP4, RRAGC were found to be mutated at relatively low frequencies, and their presence is a distinct feature of certain lymphomas or novel genetic subgroups within lymphoma types (Figure 1A). The aSHM analysis and discovery of novel hypermutated regions is handled by a custom tool Rainstorm. As a result, we were able to detect sites preferentially hypermutated in a single entity, such as the transcription start site of BACH2, mutated at lower rates than the other common target sites but significantly more in BL compared to other entities (Figure 1B). Combining aSHM at target sites discovered using our toolkit with other genetic features allowed us to explore and establish novel genetic subgroups within Burkitt lymphoma and follicular lymphoma. SV analysis can be conducted using Manta, GRIDSS, and JaBbA modules with downstream processing in GAMBLR. In B-cell lymphomas, the most common SVs identified using the automated workflow were targeting MYC, BCL2, and CCND1. Unsurprisingly, the most common translocation partner among B-cell lymphomas was the immunoglobulin heavy chain, but the novel BCL6-FOXP1, CD274-BACH2, BCL6-RHOH translocations in DLBCLs and MYC-BCL6 translocations in BLs were identified, among others (Figure 1C). Conclusions: We present here the modularized workflow for scalable and automated analysis of genomic and transcriptomic data and demonstrate that it can be successfully deployed across thousands of tumour samples for the discovery of known and novel lymphoma biology. This represents an important advancement in reproducibility that will facilitate clinical translation of genomic discoveries. Figure 1 Figure 1. Disclosures Grande: Sage Bionetworks: Current Employment. Coyle: Allakos, Inc.: Consultancy. Steidl: AbbVie: Consultancy; Trillium Therapeutics: Research Funding; Epizyme: Research Funding; Seattle Genetics: Consultancy; Curis Inc.: Consultancy; Bayer: Consultancy; Bristol-Myers Squibb: Research Funding. Scott: Abbvie: Consultancy; NanoString Technologies: Patents & Royalties: Patent describing measuring the proliferation signature in MCL using gene expression profiling.; Celgene: Consultancy; AstraZeneca: Consultancy; Incyte: Consultancy; Janssen: Consultancy, Research Funding; Rich/Genentech: Research Funding; BC Cancer: Patents & Royalties: Patent describing assigning DLBCL COO by gene expression profiling--licensed to NanoString Technologies. Patent describing measuring the proliferation signature in MCL using gene expression profiling. . Morin: Epizyme: Patents & Royalties; Celgene: Consultancy; Foundation for Burkitt Lymphoma Research: Membership on an entity's Board of Directors or advisory committees.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,006
score de la tête « metaresearch » (Gemma)0,010
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Sans objet · Signal consensuel: Sans objet
GenreSignal candidat: Logiciel · Signal consensuel: Logiciel
Score de désaccord entre enseignants0,016
Score d'incertitude au seuil0,053

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0060,010
Méta-épidémiologie (sens strict)0,0030,002
Méta-épidémiologie (sens large)0,0010,003
Bibliométrie0,0030,003
Études des sciences et des technologies0,0010,001
Communication savante0,0030,002
Science ouverte0,0030,007
Intégrité de la recherche0,0010,003
Charge utile insuffisante (le modèle a refusé de juger)0,0160,024

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,010
Tête enseignante GPT0,246
Écart entre enseignants0,236 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSans objet
Domainenon disponible
GenreLogiciel

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2021
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueBlood→Même sujetCancer Genomics and Diagnostics→Travaux en français237 207→