MétaCan
Menu
← Back to cohort
Record W3213958976 · doi:10.1182/blood-2021-147548

An Open-Source Toolkit That Powers the Genome-Wide Analysis of Mature B-Cell Lymphomas (GAMBL) Project

2021· article· en· W3213958976 on OpenAlexaff
Kostiantyn Dreval, Bruno M. Grande, Helena Winata, Jasper Wong, Lakshay Sethi, Christopher Rushton, Prasath Pararajalingam, Sarah E. Arthur, Lauren C. Chong, Brett Collinge, Krysta M. Coyle, Manuela Cruz, Stacy Hung, Shaghayegh Soudi, Nicole Thomas, Christian Steidl, David W. Scott, Ryan D. Morin, Laura K. Hilton

Bibliographic record

VenueBlood · 2021
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsCanada's Michael Smith Genome Sciences CentreSpinal Cord Injury BCSimon Fraser University
Fundersnot available
KeywordsExomeGenomeComputational biologyBiologyLymphomaGenomicsExome sequencingBioinformaticsGeneGeneticsMutation

Abstract

fetched live from OpenAlex

Abstract Introduction: Genome- and transcriptome-wide analyses continue to enhance our understanding of the molecular pathogenesis of cancer. In lymphomas, this has enabled the identification of hundreds of recurrently mutated genes, highlighting genetic heterogeneity and relationships both within and among clinical entities. While the growing availability of lymphoma genomic data sets can be leveraged to integrate genomic analyses into diagnostic testing and clinical trials, the ability to rapidly process genomic data sets in a reproducible manner serves as a barrier to this goal. To this end, we developed a suite of tools Lymphoid Cancer Research modules (LCR-modules) to facilitate the discovery of novel drivers and molecular features in lymphoma cancers and perform quantitative comparisons between disease entities. We demonstrate here how this toolkit enabled a meta-analysis of lymphoma genomic data involving genome-wide profiles of 3330 patients. Methods: We assembled a collection of whole genome, whole exome, and RNA sequencing data from a combination of controlled-access repositories and ongoing projects at BC Cancer. The scope of genomic analysis of mature B-cell lymphomas (GAMBL) project includes cell lines and patient tumors from all common mature B cell neoplasms, comprising a total of 4612 samples from 3330 patients. To facilitate the project, we developed a suite of open-source and custom bioinformatics tools (https://github.com/LCR-BCCRC/lcr-modules) that leverages the Snakemake workflow management system and includes lymphoma-centric modules for the discovery and annotation of common mutation types, analysis of B-cell receptor repertoires and discovery of novel aSHM targets and relevant non-coding mutations, and RNA-seq analysis with batch correction and normalization. Individual modules are configured to create an automated, scalable, and reproducible workflow that runs each step as dictated by the availability of new data. The cohort-level integrative analysis and comparisons across entities are handled by our custom R package GAMBLR, which facilitates open-ended data analysis and custom visualizations. Results: Simple somatic mutations (SSM) were detected using a workflow that utilizes four algorithms to identify high-confidence variants with validated default thresholds for filtering of germline variants and common FFPE-associated artifacts, allowing for processing of samples without matched normal tissue. This automated and reproducible workflow facilitated the discovery of novel genes significantly mutated across lymphomas and broadened our understanding of the scope of aberrant somatic hypermutation (aSHM) and other non-coding mutations. Specifically, HNRNPU, STAT3, TFAP4, RRAGC were found to be mutated at relatively low frequencies, and their presence is a distinct feature of certain lymphomas or novel genetic subgroups within lymphoma types (Figure 1A). The aSHM analysis and discovery of novel hypermutated regions is handled by a custom tool Rainstorm. As a result, we were able to detect sites preferentially hypermutated in a single entity, such as the transcription start site of BACH2, mutated at lower rates than the other common target sites but significantly more in BL compared to other entities (Figure 1B). Combining aSHM at target sites discovered using our toolkit with other genetic features allowed us to explore and establish novel genetic subgroups within Burkitt lymphoma and follicular lymphoma. SV analysis can be conducted using Manta, GRIDSS, and JaBbA modules with downstream processing in GAMBLR. In B-cell lymphomas, the most common SVs identified using the automated workflow were targeting MYC, BCL2, and CCND1. Unsurprisingly, the most common translocation partner among B-cell lymphomas was the immunoglobulin heavy chain, but the novel BCL6-FOXP1, CD274-BACH2, BCL6-RHOH translocations in DLBCLs and MYC-BCL6 translocations in BLs were identified, among others (Figure 1C). Conclusions: We present here the modularized workflow for scalable and automated analysis of genomic and transcriptomic data and demonstrate that it can be successfully deployed across thousands of tumour samples for the discovery of known and novel lymphoma biology. This represents an important advancement in reproducibility that will facilitate clinical translation of genomic discoveries. Figure 1 Figure 1. Disclosures Grande: Sage Bionetworks: Current Employment. Coyle: Allakos, Inc.: Consultancy. Steidl: AbbVie: Consultancy; Trillium Therapeutics: Research Funding; Epizyme: Research Funding; Seattle Genetics: Consultancy; Curis Inc.: Consultancy; Bayer: Consultancy; Bristol-Myers Squibb: Research Funding. Scott: Abbvie: Consultancy; NanoString Technologies: Patents & Royalties: Patent describing measuring the proliferation signature in MCL using gene expression profiling.; Celgene: Consultancy; AstraZeneca: Consultancy; Incyte: Consultancy; Janssen: Consultancy, Research Funding; Rich/Genentech: Research Funding; BC Cancer: Patents & Royalties: Patent describing assigning DLBCL COO by gene expression profiling--licensed to NanoString Technologies. Patent describing measuring the proliferation signature in MCL using gene expression profiling. . Morin: Epizyme: Patents & Royalties; Celgene: Consultancy; Foundation for Burkitt Lymphoma Research: Membership on an entity's Board of Directors or advisory committees.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.006
metaresearch head score (Gemma)0.010
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Software · Consensus signal: Software
Teacher disagreement score0.016
Threshold uncertainty score0.053

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0060.010
Meta-epidemiology (narrow)0.0030.002
Meta-epidemiology (broad)0.0010.003
Bibliometrics0.0030.003
Science and technology studies0.0010.001
Scholarly communication0.0030.002
Open science0.0030.007
Research integrity0.0010.003
Insufficient payload (model declined to judge)0.0160.024

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.010
GPT teacher head0.246
Teacher spread0.236 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreSoftware

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueBlood→Same topicCancer Genomics and Diagnostics→French-language works237,207→