High-grade B-cell lymphoma, not otherwise specified: an LLMPP study
Bibliographic record
Abstract
ABSTRACT: Molecular characterization of high-grade B-cell lymphoma, not otherwise specified (HGBCL-NOS), is hindered by its rarity, evolving definition, and poor diagnostic reproducibility. To address this challenge, we analyzed 92 HGBCL-NOS tumors collected across Lymphoma/Leukemia Molecular Profiling Project sites. Leveraging comparison cohorts of diffuse large B-cell lymphoma, NOS (DLBCL-NOS) and Burkitt lymphoma (BL), and molecular frameworks described in these entities, our analysis revealed a heterogenous molecular landscape, reminiscent of DLBCL-NOS but with an enrichment of BL features. By cell-of-origin classification, 59% were germinal center B-cell-like (GCB), and 25% were activated B-cell-like (ABC). LymphGen, a genetic classifier for DLBCL-NOS, assigned a genetic subtype to 34% of HGBCL-NOS. Although classification rate was lower than in DLBCL-NOS (66%), assigned subtypes spanned the spectrum of LymphGen classes, including 31% of ABCs classified as MCD. Features differentiating HGBCL-NOS from DLBCL-NOS included MYC rearrangement (47% vs 6%); dark zone signature (DZsig) expression (45% vs 7%); and more frequent mutation of ID3, MYC, CCND3, and TP53, all common to BL. A genetic classifier that differentiates DLBCL-NOS from BL classified 53% of DZsig+ tumors as BL-like, and those classified as DLBCL-like were frequently BCL2-rearranged. Among DZsig- GCB tumors, 95% were DLBCL-like. Centralized pathology review reclassified almost half of tumors as DLBCL-NOS but did not identify a more homogenous HGBCL-NOS population, with no difference in features between confirmed and reclassified tumors. In conclusion, molecular testing enables a subset of HGBCL-NOS to be assigned to established categories. Based on rarity and diagnostic challenges, broader inclusion of HGBCL-NOS should be considered in biomarker-driven DLBCL trials.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".