Artificial intelligence-aided detection for prostate cancer with multimodal routine health check-up data: an Asian multi-center study
Bibliographic record
Abstract
BACKGROUND: The early detection of high-grade prostate cancer (HGPCa) is of great importance. However, the current detection strategies result in a high rate of negative biopsies and high medical costs. In this study, the authors aimed to establish an Asian Prostate Cancer Artificial intelligence (APCA) score with no extra cost other than routine health check-ups to predict the risk of HGPCa. PATIENTS AND METHODS: A total of 7476 patients with routine health check-up data who underwent prostate biopsies from January 2008 to December 2021 in eight referral centres in Asia were screened. After data pre-processing and cleaning, 5037 patients and 117 features were analyzed. Seven AI-based algorithms were tested for feature selection and seven AI-based algorithms were tested for classification, with the best combination applied for model construction. The APAC score was established in the CH cohort and validated in a multi-centre cohort and in each validation cohort to evaluate its generalizability in different Asian regions. The performance of the models was evaluated using area under the receiver operating characteristic curve (ROC), calibration plot, and decision curve analyses. RESULTS: Eighteen features were involved in the APCA score predicting HGPCa, with some of these markers not previously used in prostate cancer diagnosis. The area under the curve (AUC) was 0.76 (95% CI:0.74-0.78) in the multi-centre validation cohort and the increment of AUC (APCA vs. PSA) was 0.16 (95% CI:0.13-0.20). The calibration plots yielded a high degree of coherence and the decision curve analysis yielded a higher net clinical benefit. Applying the APCA score could reduce unnecessary biopsies by 20.2% and 38.4%, at the risk of missing 5.0% and 10.0% of HGPCa cases in the multi-centre validation cohort, respectively. CONCLUSIONS: The APCA score based on routine health check-ups could reduce unnecessary prostate biopsies without additional examinations in Asian populations. Further prospective population-based studies are warranted to confirm these results.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".