Validation of Prognostic Models for Renal Cell Carcinoma Recurrence, Cancer-Specific Mortality, and All-Cause Mortality
Bibliographic record
Abstract
PURPOSE: Postoperative prognostic tools allow for improved prediction of future recurrence risk, patient counseling, assessment of eligibility for adjuvant treatments, and appropriate follow-up surveillance. The purpose of this analysis was to validate prognostic models for patients with kidney cancer. MATERIALS AND METHODS: The Canadian Kidney Cancer information system is a prospective cohort of patients managed at 14 institutions since January 1, 2011, to present. The Canadian Kidney Cancer information system was used to assess 15 predictive models for kidney cancer recurrence, 6 for cancer-specific mortality, and 4 for all-cause mortality in patients with a solitary, nonmetastatic kidney tumor treated with surgery (partial or radical nephrectomy). Discrimination was measured using C statistics, 5-year calibration plots for calibration, and decision curve analysis at 5 years after surgery for net benefit when considering adjuvant therapy. RESULTS: Seven thousand one hundred seventy-four patients were included. For kidney cancer recurrence, C statistics ranged from 0.62 to 0.83, depending on whether the model was derived and applied to all patients without further stratification, specific risk groups, or specific histologic subtypes. Cancer-specific mortality models had C statistics ranging from 0.60 to 0.89 and all-cause mortality models from 0.60 to 0.73. Using decision curve analysis in patients with clear-cell renal cell carcinoma, the best models for choosing adjuvant therapy to prevent recurrence and cancer-related death were the Mayo Clinic prediction models. CONCLUSIONS: Model performance varied considerably with some suitable for clinical use. If using prediction models to select adjuvant therapy, the Mayo Clinic models were best when applied to a large contemporary cohort of Canadian patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.030 | 0.063 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".