A systematic review and meta-analysis of the utility of quantitative, imaging-based approaches to predict radiation-induced toxicity in lung cancer patients
Bibliographic record
Abstract
BACKGROUND AND PURPOSE: To conduct a systematic review and meta-analysis of the performance of radiomics, dosiomics and machine learning in generating toxicity prediction in thoracic radiotherapy. MATERIALS AND METHODS: An electronic database search was conducted and dual-screened by independent authors to identify eligible studies for systematic review and meta-analysis. Data was extracted and study quality was assessed using TRIPOD for machine learning studies, RQS for Radiomics and RoB for dosiomics. RESULTS: 10,703 studies were identified, and 5,252 entered screening. 104 studies including 23,373 patients were eligible for systematic review. Primary toxicity predicted was radiation pneumonitis (81), followed by esophagitis (12) and lymphopenia (4). Fourty-two studies studying radiation pneumonitis were eligible for meta-analysis, with pooled area-under-curve (AUC) of 0.82 (95% CI 0.79-0.85). Studies with machine learning had the best performance, with classical and deep learning models having similar performance. There is a trend towards an improvement of the performance of models with the year of publication. There is variability in study quality among the three study categories and dosiomic studies scored the highest among these. Publication bias was not observed. CONCLUSION: The majority of existing literature using radiomics, dosiomics and machine learning has focused on radiation pneumonitis prediction. Future research should focus on toxicity prediction of other organs at risk and the adoption of these models into clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.008 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".