Evaluating Outcome Prediction via Baseline, End-of-Treatment, and Delta Radiomics on PET-CT Images of Primary Mediastinal Large B-Cell Lymphoma
Bibliographic record
Abstract
Objectives: Accurate outcome prediction is important for making informed clinical decisions in cancer treatment. In this study, we assessed the feasibility of using changes in radiomic features over time (Delta radiomics: absolute and relative) following chemotherapy, to predict relapse/progression and time to progression (TTP) of primary mediastinal large B-cell lymphoma (PMBCL) patients. Material and Methods: Given the lack of standard staging PET scans until 2011, only 31 out of 103 PMBCL patients in our retrospective study had both pre-treatment and end-of-treatment (EoT) scans. Consequently, our radiomics analysis focused on these 31 patients who underwent [18F]FDG PET-CT scans before and after R-CHOP chemotherapy. Expert manual lesion segmentation was conducted on their scans for delta radiomics analysis, along with an additional 19 EoT scans, totaling 50 segmented scans for single time point analysis. Radiomics features (on PET and CT), along with maximum and mean standardized uptake values (SUVmax and SUVmean), total metabolic tumor volume (TMTV), tumor dissemination (Dmax), total lesion glycolysis (TLG), and the area under the curve of cumulative standardized uptake value-volume histogram (AUC-CSH) were calculated. We additionally applied longitudinal analysis using radial mean intensity (RIM) changes. For prediction of relapse/progression, we utilized the individual coefficient approximation for risk estimation (ICARE) and machine learning (ML) techniques (K-Nearest Neighbor (KNN), Linear Discriminant Analysis (LDA), and Random Forest (RF)) including sequential feature selection (SFS) following correlation analysis for feature selection. For TTP, ICARE and CoxNet approaches were utilized. In all models, we used nested cross-validation (CV) (with 10 outer folds and 5 repetitions, along with 5 inner folds and 20 repetitions) after balancing the dataset using Synthetic Minority Oversampling TEchnique (SMOTE). Results: To predict relapse/progression using Delta radiomics between the baseline (staging) and EoT scans, the best performances in terms of accuracy and F1 score (F1 score is the harmonic mean of precision and recall, where precision is the ratio of true positives to the sum of true positives and false positives, and recall is the ratio of true positives to the sum of true positives and false negatives) were achieved with ICARE (accuracy = 0.81 ± 0.15, F1 = 0.77 ± 0.18), RF (accuracy = 0.89 ± 0.04, F1 = 0.87 ± 0.04), and LDA (accuracy = 0.89 ± 0.03, F1 = 0.89 ± 0.03), that are higher compared to the predictive power achieved by using only EoT radiomics features. For the second category of our analysis, TTP prediction, the best performer was CoxNet (LASSO feature selection) with c-index = 0.67 ± 0.06 when using baseline + Delta features (inclusion of both baseline and Delta features). The TTP results via Delta radiomics were comparable to the use of radiomics features extracted from EoT scans for TTP analysis (c-index = 0.68 ± 0.09) using CoxNet (with SFS). The performance of Deauville Score (DS) for TTP was c-index = 0.66 ± 0.09 for n = 50 and 0.67 ± 03 for n = 31 cases when using EoT scans with no significant differences compared to the radiomics signature from either EoT scans or baseline + Delta features (p-value> 0.05). Conclusion: This work demonstrates the potential of Delta radiomics and the importance of using EoT scans to predict progression and TTP from PMBCL [18F]FDG PET-CT scans.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".