The New Prostate Cancer Grading System Does Not Improve Prediction of Clinical Recurrence After Radical Prostatectomy: Results of a Large, Two‐Center Validation Study
Bibliographic record
Abstract
BACKGROUND: A new prostate cancer (PCa) grading system (namely, Gleason score-GS- ≤6 vs. 3 + 4 vs. 4 + 3 vs. 8 vs. ≥9) was recently proposed and assessed on biochemical recurrence (BCR) showing improved predictive abilities compared to the commonly used three-tier system (GS ≤6 vs. 7 vs. ≥8). We assessed the predictive ability of the five-tier grade group (GG) system on harder clinical endpoint, namely clinical recurrence (CR). METHODS: Between 2005 and 2014, 9,728 clinically localized PCa patients were treated with radical prostatectomy (RP) at two tertiary referral centers. Kaplan-Meier curves, multivariable Cox regression analyses, and concordance index (C-index) were used to assess CR after treatment according to four Gleason grade classifications at biopsy and RP: Group 1: ≤6 versus 7 versus ≥8; Group 2: ≤6 versus 3 + 4 vs. 4 + 3 versus ≥8; Group 3: ≤6 versus 7 versus 8 versus ≥9; Group 4: ≤6 versus 3 + 4 versus 4 + 3 versus 8 versus ≥9. Same analyses were repeated in patients who had BCR (n = 1,624). Decision curve analyses were performed to evaluate and compare the net benefit associated with the use of the four Gleason grade classifications. RESULTS: Overall, 443 (4.6%) patients had CR. The hazard ratio of the GS 3 + 4, 4 + 3, 8, and ≥9 relative to GS ≤6 were 3.63, 5.93, 11.44, 18.08 and 4.93, 9.99, 15.31 and 25.12 in the pre- and post-treatment models, respectively. The C-index of the five-tier GG system was slightly higher relative to the other 3 Gleason grade classifications both in the pre- (range: 0.001-0.006) and post-treatment models (range: 0-0.008). Similar findings were observed when we focused our analyses in patients with BCR after RP. The use of the five-tier GG system did not result into higher net-benefit relative to the other three Gleason grade classifications. CONCLUSIONS: The difference in accuracy between the five-tier GG system and the other Gleason grade classifications, using CR as an endpoint, is clinically negligible. Current evidence suggests that the five-tier GG system represents a simplified user-friendly scheme available for patient counseling rather than a new histopathological diagnostic system that improves the prediction of CR. Prostate 77:263-273, 2017. © 2016 Wiley Periodicals, Inc.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".