A Systematic Review and Meta-Analysis of Prognostic Nomograms After UTUC Surgery
Bibliographic record
Abstract
Background: Current guidelines recommend assessing the prognosis in high-risk upper tract urothelial carcinoma patients (UTUC) after surgery. However, no specific method is endorsed. Among the various prognostic models, nomograms represent an easy and accurate tool to predict the individual probability for a specific event. Therefore, identifying the best-suited nomogram for each setting seems of great interest to the patient and provider. Objectives: To identify, summarize and compare postoperative UTUC nomograms predicting oncologic outcomes. To estimate the overall performance of the nomograms and identify the most reliable predictors. To create a reference tool for postoperative UTUC nomograms, physicians can use in clinical practice. Design: A systematic review was conducted following the recommendations of Cochrane's Prognosis Methods Group. Medline and EMBASE databases were searched for studies published before December 2021. Nomograms were grouped according to outcome measurements, the purpose of use, and inclusion and exclusion criteria. Random-effects meta-analyses were performed to estimate nomogram group performance and predictor reliability. Reference tables summarizing the nomograms' important characteristics were created. Results: The systematic review identified 26 nomograms. Only four were externally validated. Study heterogeneity was significant, and the overall Risk of Bias (RoB) was high. Nomogram groups predicting overall survival (OS), recurrence-free survival (RFS), and intravesical recurrence (IVR) had moderate discrimination accuracy (c-Index summary estimate with 95% confidence interval [95% CI] and prediction interval [PI] > 0.6). Nomogram groups predicting cancer-specific survival (CSS) had good discrimination accuracy (c-Index summary estimate with 95% CI and PI > 0.7). Advanced pathological tumor stage (≥ pT3) was the most reliable predictor of OS. Pathological tumor stage (≥ pT2), age, and lymphovascular invasion (LVI) were the most reliable predictors of CSS. LVI was the most reliable predictor of RFS. Conclusions: Despite a moderate to good discrimination accuracy, severe heterogeneity discourages the uninformed use of postoperative prognostic UTUC nomograms. For nomograms to become of value in a generalizable population, future research must invest in external validation and assessment of clinical utility. Meanwhile, this systematic review serves as a reference tool for physicians choosing nomograms based on individual needs. Systematic Review Registration: https://www.crd.york.ac.uk/prospero/display_record.php?RecordID=282596, identifier PROSPERO [CRD42021282596].
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.020 | 0.003 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".