External Validation and Comparison of Prostate Cancer Risk Calculators Incorporating Multiparametric Magnetic Resonance Imaging for Prediction of Clinically Significant Prostate Cancer
Bibliographic record
Abstract
PURPOSE: We sought to externally validate recently published prostate cancer risk calculators incorporating multiparametric magnetic resonance imaging to predict clinically significant prostate cancer. We also compared the performance of these calculators to that of multiparametric magnetic resonance imaging naïve prostate cancer risk calculators. MATERIALS AND METHODS: We identified men without a previous prostate cancer diagnosis who underwent transperineal template saturation prostate biopsy with fusion guided targeted biopsy between November 2014 and March 2018 at our academic tertiary referral center. Any Gleason pattern 4 or greater was defined as clinically significant prostate cancer. Predictors, which were patient age, prostate specific antigen, digital rectal examination, prostate volume, family history, previous prostate biopsy and the highest region of interest according to the PI-RADS™ (Prostate Imaging Reporting and Data System), were retrospectively collected. Four multiparametric magnetic resonance imaging prostate cancer risk calculators and 2 multiparametric magnetic resonance imaging naïve prostate cancer risk calculators were evaluated for discrimination, calibration and the clinical net benefit using ROC analysis, calibration plots and decision curve analysis. RESULTS: Of the 468 men 193 (41%) were diagnosed with clinically significant prostate cancer. Three multiparametric magnetic resonance imaging prostate cancer risk calculators showed similar discrimination with a ROC AUC significantly higher than that of the other prostate cancer risk calculators (AUC 0.83-0.85 vs 0.69-0.74). Calibration in the large showed 2% deviation from the true amount of clinically significant prostate cancer for 2 multiparametric magnetic resonance imaging risk calculators while the other calculators showed worse calibration at 11% to 27%. A clinical net benefit was observed only for 3 multiparametric magnetic resonance imaging risk calculators at biopsy thresholds of 15% or greater. None of the 6 investigated prostate cancer risk calculators demonstrated clinical usefulness against a biopsy all strategy at thresholds less than 15%. CONCLUSIONS: The performance of multiparametric magnetic resonance imaging prostate cancer risk calculators varies but they generally outperform multiparametric magnetic resonance imaging naïve prostate cancer risk calculators in regard to discrimination, calibration and clinical usefulness. External validation in other biopsy settings is highly encouraged.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.043 | 0.109 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".