Models and Biomarkers for Local Response Prediction in Early-Stage and Oligometastatic Non-small Cell Lung Cancer Patients Treated With Stereotactic Body Radiation Therapy Using Machine Learning
Bibliographic record
Abstract
Background A minority of patients receiving stereotactic body radiation therapy (SBRT) for non-small cell lung cancer (NSCLC) are not good responders. Radiomic features can be used to generate predictive algorithms and biomarkers that can determine treatment outcomes and stratify patients to their therapeutic options. This study investigated and attempted to validate the radiomic and clinical features obtained from early-stage and oligometastatic NSCLC patients who underwent SBRT, to predict local response. Methodology A single-institution, Institutional Review Board (IRB)-approved retrospective review was conducted on adult patients with early-stage and oligometastatic SBRT-treated NSCLC at the Jewish General Hospital. The study included 98 patients (82 with early-stage NSCLC and 16 with oligometastatic disease), with a median age of 76 years and a male-to-female ratio of 46:52. A total of 116 lesions were treated with SBRT between 2009 and 2022. Radiomics features (n = 107) were extracted from CT planning scans using PyRadiomics, and clinical data were collected for all 98 patients. Local response was assessed according to Response Evaluation Criteria In Solid Tumors (RECIST 1.1) criteria. Classification models, including support vector machines, random forests, adaptive boosting, and multi-layer perceptrons (MLPs), were used. Models were trained using a fivefold cross-validation scheme. Their performances were measured with receiver operating characteristic plots on the validation folds. Using the importance of the permutation feature, predictive biomarkers were identified. Results The most predictive model, incorporating all patients and using an MLP classifier with Adaptive Synthetic (ADASYN) sampling, a combined-input approach, and a radiomic filter, achieved an area under the curve (AUC) of 0.94 ± 0.05. When oligometastatic patients were omitted, the best model (AUC 0.95 ± 0.06) was also predictive, using a support vector classification (SVC) radial basis function (RBF) classifier, ADASYN sampling, and a clinical-based input. Treatment site and performance status, along with radiomic features such as first-order root-mean-squared-intensity, first-order skewness, and gray-level nonuniformity, were found to be predictive biomarkers. Conclusions The predictive models generated and the biomarkers identified could be used in clinical decision support systems for SBRT-treated NSCLC patients. Additionally, treatment site, performance status, and radiomic features were the most predictive variables.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".