Avoiding Unnecessary Biopsy: MRI-based Risk Models versus a PI-RADS and PSA Density Strategy for Clinically Significant Prostate Cancer
Bibliographic record
Abstract
Background In validation studies, risk models for clinically significant prostate cancer (csPCa; Gleason score ≥3+4) combining multiparametric MRI and clinical factors have demonstrated poor calibration (over- and underprediction) and limited use in avoiding unnecessary prostate biopsies. Purpose MRI-based risk models following local recalibration were compared with a strategy that combined Prostate Imaging Data and Reporting System (PI-RADS; version 2) and prostate-specific antigen density (PSAd) to assess the potential reduction of unnecessary prostate biopsies. Materials and Methods This retrospective study included 385 patients without prostate cancer diagnosis who underwent multipara-metric MRI (PI-RADS category ≥3) and MRI-targeted biopsy between 2015 and 2019. Recalibration and selection of the best-performing MRI model (MRI–European Randomized Study of Screening for Prostate Cancer [ERSPC], van Leeuwen, Radtke, and Mehralivand models) were undertaken in cohort C1 (n = 242; 2015–2017). The impact on biopsy decisions was compared with an alternative strategy (no biopsy for PI-RADS category 3 plus PSAd < 0.1 ng/mL per milliliter) in cohort C2 (n = 143; 2018–2019). Discrimination, calibration, and clinical utility were assessed by using the area under the receiver operating characteristic curve (AUC), calibration plots, and decision curve analysis, respectively. Results The prevalence of csPCa was 38% (93 of 242 patients) and 45% (64 of 143 patients) in cohorts C1 and C2, respectively. Decision curve analysis demonstrated the highest net benefit for the van Leeuwen and Mehralivand models in C1. Used for biopsy decisions in C2, van Leeuwen (AUC, 0.84; 95% CI: 0.77, 0.9) and Mehralivand (AUC, 0.79; 95% CI: 0.72, 0.86) enabled no net benefit at a risk threshold of 10%. Up to a risk threshold of 15%, net benefit remained inferior to the PI-RADS plus PSAd strategy, which avoided biopsy in 63 per 1000 men, without missing csPCa. Without prior recalibration in C1, three of four models (MRIERSPC, Radtke, Mehralivand) were poorly calibrated and not clinically useful in C2. Conclusion The number of unnecessary prostate biopsies in men with positive MRI may be safely reduced by using a prostate-specific antigen density–based strategy. In a risk-averse scenario, this strategy enabled better biopsy decisions compared with MRI-based risk models. ©RSNA, 2021 Online supplemental material is available for this article.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".