Genomic and Transcriptional Characterization of Primary Mediastinal Large B Cell Lymphoma
Bibliographic record
Abstract
Abstract Introduction: Primary mediastinal large B-cell lymphoma (PMBL) is a rare non-Hodgkin lymphoma subtype that occurs predominantly in young adults, with an overall favorable prognosis. The cell of origin is presumed to be thymic medullary B-cells and the gene expression profile of PMBL is similar to classic Hodgkin lymphoma. Recent studies have begun unravelling the genomic alterations underlying PMBL. Frequent, recurrent mutations (e.g. B2M, TNFAIP3, SOCS1, STAT6, GNA13) have been reported, but most of the studies have analyzed a small number of cases. To gain further insights into disease biology, we recruited 63 cases of PMBL as part of the Atlas of Blood Cancer Genomes (ABC-G) initiative, a consortium consisting of 25 institutions. Methods: Formalin-fixed paraffin-embedded (FFPE) biopsies and clinical data were collected. All cases were subjected to centralized review by an experienced panel of hematopathologists to ensure accurate diagnosis. Whole-exome DNA and RNA sequencing was performed using the Illumina platform and the DNA and RNA reads aligned to the GRCh38 genome and transcriptome respectively. Exonic variants were filtered using a set of paired normal samples and population-based databases to identify putative driver mutations, which were then aggregated at the gene level. Mutational analysis was performed on 56 samples that passed quality filtering and expression analysis on 45 samples. RNAseq data was normalized using DESeq2. Results: The cohort included samples from 16 males and 24 females, with a median age of 33 years (range 16 - 72) at the time of diagnosis. The majority of patients were treated with R-CHOP (47%) or R-EPOCH (43%), with 93% of patients surviving through the end of follow-up (median follow-up: 60.1 months). Besides the known recurrent mutations involving the JAK-STAT (STAT6 -21%, SOCS1 - 26%), NFKB (TNFAIP3 - 27%, NFKB1A - 7%), immune escape (B2M - 20%, LTB - 11%, IRF8 - 9%, IRF4 -9%), and chromatin modification (ZNF217 - 16%, CREBBP - 11%, KMT2D -11%) pathways , we discovered recurrent somatic variants in novel candidate driver genes in this disease, including NOTCH4 (7%), DICER1 (11%), MCL1 (7%), amongst others. EZH2, EP300, and XPO1 mutations were not detected. CIITA mutations and fusions were observed in 14% and 11% of cases, respectively, with novel partner genes (IGHA2, IGHG1, CDC6) detected in 67% of the fusion positive cases. Copy number alterations included gains at 2p16.1 (REL - 20%) and 9p24.2 (JAK2/PDL1/PDL2 - 24%), as well as loci not previously implicated in PMBL, 8q24.3 and 9q34.3 (each in 20%). Of note, CIITA alterations and 9p24 gains were virtually mutually exclusive, highlighting diverse mechanisms of immune escape in this entity. The transcriptomes of cases harboring CIITA alterations demonstrated differential enrichment of genes involved in protein glycosylation. The PMBLs in our series showed significant enrichment of the reported PMBL genetic classifier score, compared to nodal diffuse large B cell lymphoma (DLBCL) (p=0.0003). Finally, the gene expression profile of thymic B cells was more similar to that of PMBL than nodal DLBCL (p=0.0144). Conclusions: Our study, representing one of the largest comprehensive genomic and transcriptomic analyses of PMBL, expands the mutational landscape of PMBL, provides evidence for biologically distinct disease subsets and suggests an origin of PMBLs from thymic B-cells. Disclosures Hsi: AbbVie: Research Funding; Eli Lilly: Research Funding; Cytomx: Honoraria; Seattle Genetics: Honoraria. McKinney: BTG: Consultancy; Celgene: Consultancy, Research Funding; Epizyme: Consultancy; Genetech: Consultancy, Honoraria, Research Funding; Incyte: Research Funding; Kite/Gilead: Honoraria, Speakers Bureau; Molecular Templates: Consultancy, Research Funding; Nordic Nanovector: Research Funding; Novartis: Research Funding; Pharmacyclics: Consultancy; Verastem: Consultancy; Beigene: Research Funding; ADC Therapeutics: Consultancy, Speakers Bureau. Jaye: Stemline Therapeutics: Honoraria. Cohen: Genentech, Takeda, BMS/Celgene, BioInvent, LAM, Astra Zeneca, Novartis, Loxo/Lilly: Research Funding; Janssen, Adaptive, Aptitude Health, BeiGene, Cellectar, Adicet, Loxo/Lilly, AStra ZenecaKite/Gilead: Consultancy. Behdad: Lilly: Speakers Bureau; Roche/Foundation Medicine: Speakers Bureau; Thermo Fisher: Speakers Bureau. Dave: Data Driven Bioscience: Current equity holder in publicly-traded company.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".