56 | GENOMIC ANALYSIS OF MATURE B‐CELL LYMPHOMAS
Bibliographic record
Abstract
K. Dreval, M. Cruz, and L. K. Hilton equally contributing authors. Introduction: The B-cell lymphomas represent a genetically heterogeneous and complex collection of malignancies. Among diffuse large B-cell lymphoma (DLBCL), Burkitt lymphoma (BL) and follicular lymphoma (FL), over 140 recurrently mutated genes have been established. Subdivisions of these pathologies into molecular or genetic subgroups with distinct biological features is of interest but it remains unclear whether genome-wide analyses have identified a sufficient number of relevant genetic features. To search for additional drivers and refine our understanding of common mutation patterns across the mature B-cell lymphomas, we performed a comprehensive meta-analysis of all available published and locally generated whole genome sequencing (WGS) and exome data. Methods: Sequencing data was assembled from a total of 2603 DLBCL, 784 FL, 433 BL, 202 MCL, 213 CLL and 808 cases spanning other mature B-cell lymphoma pathologies. This includes WGS and exome data from 1992 and 3051 samples, respectively. All WGS and exome data was analyzed for somatic mutations and structural variants using LCR-modules, our suite of open-source pipelines. RNA-seq data, available from 1837 cases, was analyzed for gene expression, alternative splicing and detecting oncogene rearrangements. Significantly-mutated genes (SMGs) were identified using a combination of MutSigCV, OncodriveFML and dNdSCV. Non-coding loci enriched for mutations were comprehensively identified using FishHook. Results: Using a pooled analysis of all DLBCL, BL and FL samples with paired normals, we identified 133 SMGs. Of these genes, 106 were among previously reported high confidence SMGs, with the remaining 27 not previously attributed to these pathologies. The mutation incidence among these new genes was low (median: 2.46), as expected. Notable examples are genes with potential roles in chromatin remodeling (ARID1B, INO80), immune evasion (FCGR2B), DNA damage response (RBM38), and BCR signaling (CD79A). While most of the novel genes were more commonly mutated in DLBCL, CDKN2C and FIP1L1 mutations were more abundant in BL. Despite the volume of data, this analysis did not reproduce 37 of the genes that have previously been attributed to at least one of these entities. Most of these represent targets of aSHM, such as BTG1, CIITA and ACTG1, which may be enriched for passenger mutations. Recurrence analysis identified a total of 105 mutation hotspots in 66 genes. This also recovered additional genes with significant hotspots that were not globally significant, including TLR2, BCOR, BCR and MEF2C. We found 130 non-coding loci that were enriched for mutations, with most regions having the highest mutation burden in DLBCL. Conclusions: Genomic Analysis of Mature B-cell Lymphomas (GAMBL) analysis has extended the list of lymphoma genes to 170 and has revealed the existence of mutation hotspots in more than a third of these genes and many non-coding loci with regulatory potential. Keywords: bioinformatics; computational and systems biology; genomics, epigenomics, and other -omics; aggressive B-cell non-Hodgkin lymphoma No potential sources of conflict of interest.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".