Computational pathology to predict docetaxel benefit for high-risk localized prostate cancer in NRG/RTOG 0521 (NCT00288080).
Bibliographic record
Abstract
1557 Background: The benefit of adding docetaxel (DTX) to standard of care (SOC) for high-risk localized prostate cancer remains debated. The NRG/RTOG 0521 randomized phase III trial demonstrated that docetaxel, when added to SOC—comprising radiotherapy (RT) and long-term androgen deprivation therapy (ADT)—improved overall survival (OS). However, while RTOG 0521 demonstrated improved OS with DTX, the observed improvement did not meet the predetermined threshold for clinical significance, leaving the role of DTX intensification uncertain. Enhanced stratification methods are needed to identify aggressive disease phenotypes and guide patient selection for adjuvant chemotherapy. This study aims to develop and validate a computational AI derived pathology image classifier (APIC) to quantify the tumor-immune microenvironment from diagnostic biopsy specimens and predict DTX benefit in patients from the NRG/RTOG 0521 trial. Methods: The study included patients with available high-quality biopsy images from the NRG/RTOG 0521 trial. Primary outcome was OS, median follow-up was 5.7 years. After segmenting nuclei and identifying lymphocytes, we derived features that captured immune-tumor spatial patterns and nuclear diversity in the tumor microenvironment to construct APIC. DTX benefit was evaluated using Cox proportional hazards models with interaction terms, log-rank tests and Kaplan-Meier analyses by comparing OS between treatment arms within APIC-stratified groups. Results: Among NRG/RTOG 0521 trial participants, 350 patients had evaluable quality biopsy slide images. Half of the SOC (RT+ADT) arm was used for training (84 patients), and 266 patients were used for validation (SOC: 85 patients, and SOC+DTX arm: 181 patients). DTX significantly improved OS in APIC-positive (n = 119, 45%) patients (HR = 0.49, 95% CI: 0.26-0.92, p = 0.023) but not in APIC-negative (n = 147, 55%) patients (HR = 1.17, 95% CI: 0.59-2.3, p = 0.66). APIC-positive patients derived 22% 10-year OS benefit (95% CI: 1.7%-41.6%) from DTX. The 10-year OS was 74% in the DTX arm compared to 52% with RT and ADT alone in the APIC-positive group. A significant interaction (p = 0.024) was observed between APIC status and treatment. Conclusions: We validated APIC as a predictive biomarker for DTX benefit in high-risk localized prostate cancer patients from NRG/RTOG 0521, identifying a subset who achieved significant survival improvement from treatment intensification – a benefit not reached in the unselected trial population. Further investigation is warranted to evaluate APIC's predictive potential of DTX intensification in metastatic disease settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".