Data from A New Immunostain Algorithm Classifies Diffuse Large B-Cell Lymphoma into Molecular Subtypes with High Accuracy
Bibliographic record
Abstract
Abstract Purpose: Hans and coworkers previously developed an immunohistochemical algorithm with ∼80% concordance with the gene expression profiling (GEP) classification of diffuse large B-cell lymphoma (DLBCL) into the germinal center B-cell–like (GCB) and activated B-cell–like (ABC) subtypes. Since then, new antibodies specific to germinal center B-cells have been developed, which might improve the performance of an immunostain algorithm. Experimental Design: We studied 84 cases of cyclophosphamide-doxorubicin-vincristine-prednisone (CHOP)–treated DLBCL (47 GCB, 37 ABC) with GCET1, CD10, BCL6, MUM1, FOXP1, BCL2, MTA3, and cyclin D2 immunostains, and compared different combinations of the immunostaining results with the GEP classification. A perturbation analysis was also applied to eliminate the possible effects of interobserver or intraobserver variations. A separate set of 63 DLBCL cases treated with rituximab plus CHOP (37 GCB, 26 ABC) was used to validate the new algorithm. Results: A new algorithm using GCET1, CD10, BCL6, MUM1, and FOXP1 was derived that closely approximated the GEP classification with 93% concordance. Perturbation analysis indicated that the algorithm was robust within the range of observer variance. The new algorithm predicted 3-year overall survival of the validation set [GCB (87%) versus ABC (44%); P < 0.001], simulating the predictive power of the GEP classification. For a group of seven primary mediastinal large B-cell lymphoma, the new algorithm is a better prognostic classifier (all “GCB”) than the Hans' algorithm (two GCB, five non-GCB). Conclusion: Our new algorithm is significantly more accurate than the Hans' algorithm and will facilitate risk stratification of DLBCL patients and future DLBCL research using archival materials. (Clin Cancer Res 2009;15(17):5494–502)
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".