ESUR/ESUI position paper: developing artificial intelligence for precision diagnosis of prostate cancer using magnetic resonance imaging
Bibliographic record
Abstract
Artificial intelligence developments are essential to the successful deployment of community-wide, MRI-driven prostate cancer diagnosis. AI systems should ensure that the main benefits of biopsy avoidance are delivered while maintaining consistent high specificities, at a range of disease prevalences. Since all current artificial intelligence / computer-aided detection systems for prostate cancer detection are experimental, multiple developmental efforts are still needed to bring the vision to fruition. Initial work needs to focus on developing systems as diagnostic supporting aids so their results can be integrated into the radiologists' workflow including gland and target outlining tasks for fusion biopsies. Developing AI systems as clinical decision-making tools will require greater efforts. The latter encompass larger multicentric, multivendor datasets where the different needs of patients stratified by diagnostic settings, disease prevalence, patient preference, and clinical setting are considered. AI-based, robust, standard operating procedures will increase the confidence of patients and payers, thus enabling the wider adoption of the MRI-directed approach for prostate cancer diagnosis. KEY POINTS: • AI systems need to ensure that the benefits of biopsy avoidance are delivered with consistent high specificities, at a range of disease prevalence. • Initial work has focused on developing systems as diagnostic supporting aids for outlining tasks, so they can be integrated into the radiologists' workflow to support MRI-directed biopsies. • Decision support tools require a larger body of work including multicentric, multivendor studies where the clinical needs, disease prevalence, patient preferences, and clinical setting are additionally defined.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.015 | 0.017 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.004 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.006 | 0.006 |
| Insufficient payload (model declined to judge) | 0.025 | 0.017 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".