CD5 Gene Signature Identifies Diffuse Large B-Cell Lymphomas Sensitive to Brutonʼs Tyrosine Kinase Inhibition
Bibliographic record
Abstract
PURPOSE: A genetic classifier termed LymphGen accurately identifies diffuse large B-cell lymphoma (DLBCL) subtypes vulnerable to Bruton's tyrosine kinase inhibitors (BTKis), but is challenging to implement in the clinic and fails to capture all DLBCLs that benefit from BTKi-based therapy. Here, we developed a novel CD5 gene expression signature as a biomarker of response to BTKi-based therapy in DLBCL. METHODS: CD5 immunohistochemistry (IHC) was performed on 404 DLBCLs to identify CD5 IHC+ and CD5 IHC- cases, which were subsequently characterized at the molecular level through mutational and transcriptional analyses. A 60-gene CD5 gene expression signature (CD5sig) was constructed using genes differentially expressed between CD5 IHC+ and CD5 IHC- non-germinal center B-cell-like (non-GCB DLBCL) DLBCLs. This CD5sig was applied to external DLBCL data sets, including pretreatment biopsies from patients enrolled in the PHOENIX study (n = 584) to define the extent to which the CD5sig could identify non-GCB DLBCLs that benefited from the addition of ibrutinib to rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone (R-CHOP). RESULTS: DLBCLs lacked canonical BCR-activating mutations or were LymphGen-unclassifiable (LymphGen-Other). The CD5sig recapitulated these findings in multiple independent data sets, indicating its utility in identifying DLBCLs with genetic and nongenetic bases for BCR dependence. Supporting this notion, CD5sig+ DLBCLs derived a selective survival advantage from the addition of ibrutinib to R-CHOP in the PHOENIX study, independent of LymphGen classification. CONCLUSION: CD5sig is a useful biomarker to identify DLBCLs vulnerable to BTKi-based therapies and complements current biomarker approaches by identifying DLBCLs with genetic and nongenetic bases for BTKi sensitivity.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".