MétaCan
Menu
← Back to cohort
Record W4389248467 · doi:10.1182/blood-2023-181710

A New Machine Learning-Based Prognostic Prediction Strategy for Patients with Diffuse-Large B-Cell Lymphoma

2023· article· en· W4389248467 on OpenAlexaff
Valentin Daguerre, Élisabeth Daguenet, Nicolas Jacquet-Francillon, Ludovic Fouillet, Ugo Thévenet, Maxime Bonjour, Hervé Ghesquières, Imran Ahmad, Jérôme Cornillon

Bibliographic record

VenueBlood · 2023
Typearticle
Languageen
FieldMedicine
TopicCAR-T cell therapy research
Canadian institutionsUniversité de MontréalHôpital Maisonneuve-Rosemont
Fundersnot available
KeywordsDiffuse large B-cell lymphomaMedicineInternal medicineOncologyLymphomaCHOPABVDRituximabChemotherapyCyclophosphamideVincristine

Abstract

fetched live from OpenAlex

Introduction - Diffuse large B-cell lymphoma (DLBCL) is a frequent and aggressive lymphoma. First line of treatment associates an anti-CD20 immunotherapy with an anthracycline-based chemotherapy (R-CHOP). After this treatment, 30% of patients present primary refractory disease or early relapse within two years (R/R DLBCL). At date, CAR-T cells improve the prognosis of patients with R/R DLBCL compared with high-dose chemotherapy followed by autologous stem-cell transplant. However, predicting this failure remains a challenge at time of diagnosis. Current standard tools as the R-IPI score have many weaknesses, particularly for patients most at risk. A multitude of clinical, biological and PET data are available and could help to precise R/R DLBCL probability. In this study, we propose a new strategy based on machine learning and exploitation of a large dataset to early detect patients with high risk of R/R DLBCL at diagnosis. Material and methods - Between August 2012 and November 2021, we retrospectively collected a cohort of 218 patients with DLBCL receiving first-line treatment R-CHOP. Twenty-two parameters were evaluated for each patient, including well-known predictive factors such as R-IPI components, but also advanced metabolic parameters (SUVmax, total metabolic tumor volume (TMTV), total lesion glycolysis (TLG) and Dmax), as well as pathological and biological features. We used supervised machine learning (ML) to build a multidimensional model to predict POD24, which is defined by primary refractory disease or relapse within two years. Conventional binomial logistic regression was also performed to create a conventional predictive model. In parallel, a group of three hematologists was asked to generate a consensus about POD24 occurrence using Delphi method. Results - Median age was 70 years. With a median follow up of 46 months, complete response rate was 87% and overall response rate was 89%. Overall survival and progression-free survival were respectively 81.6 % (95%CI 76.6 - 86.9) and 70.6 % (95%CI 64.8 - 76.9) at 24 months. POD24 event occurred in 23,5 % of patients. Performance status, Ann-Arbor staging, number of extranodal sites, medullar infiltration, TMTV, TLG, Dmax, lymphocytes and LDH showed significant predictability of POD24 when considered in univariate analysis. We used forward selection of variables to build a conventional predictive model including LDH, lymphocytes and Dmax. Area under the ROC curve was 0,72. Secondly, a ML model was trained on 147 patients and was applied on a test cohort with 49 remaining patients. Prediction accuracy was 89.7% and area under the ROC curve was 0,91. Decision threshold was raised to 0.75 cut-off, allowing to predict 60% of POD24 with 100% specificity. The four predominant variables in the model were Dmax, hepatic SUV, LDH and monocytes. Additionally, experts were confronted to 55 patients randomly sampled from the cohort, 14 (25%) of whom showed POD24. Prediction accuracy was 71% and area under the ROC curve was 0,73. Seventy-nine percents of POD24 was predicted by experts, with 68% specificity. Conclusion - We present a ML model that was built to predict POD24 and then R/R situations in DLBCL patients. It showed great accuracy with high specificity rate. In new immunotherapies and CAR-T cells era, this proof-of-concept study highlights the fact that ML models could be a helpful to decide on first-line strategies for DLBCL patients. Larger multicentric study is necessary to confirm and validate this model for clinical use.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.003
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.003
Threshold uncertainty score0.006

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.003
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.017
GPT teacher head0.263
Teacher spread0.246 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueBlood→Same topicCAR-T cell therapy research→French-language works237,207→