A non-invasive MRI-based multimodal fusion deep learning model (MF-DLM) for predicting overall survival in bladder cancer: a multicentre retrospective study
Bibliographic record
Abstract
Background: Accurate prognosis prediction in bladder cancer (BCa) is crucial for personalized treatment. This study aimed to develop and validate a non-invasive model using magnetic resonance imaging (MRI) for predicting the overall survival (OS) in patients with BCa. Methods: This retrospective multicentre study included 1131 patients with BCa from eight institutions in China from June 2011 to March 2024. 871 patients were enrolled from one centre, who were randomly divided (8:2) into training (n = 697) and internal validation (n = 174) sets. For the external test set, 260 patients with BCa from seven centres were retrospectively included. We developed a multimodal fusion deep learning model (MF-DLM), leveraging a cross-attention mechanism to integrate four key preoperative data modalities: three-dimensional (3D) deep learning features using a modified 3D ResNet50 network, 3D radiomics features, morphological MRI features, and clinical features. Patients were stratified into low- and high-risk prognostic groups based on MF-DLM scores, and model interpretability was evaluated using Shapley additive explanations (SHAP) and Gradient-weighted class activation mapping (Grad-CAM). Findings: The median follow-up time for the training, validation, and external test sets are 38.0 months (interquartile ranges [IQR]: 22.0, 62.0), 40.5 months (IQR: 23.0, 71.0), and 38.5 months (IQR: 26.0, 50.0), respectively. The MF-DLM demonstrated excellent performance in predicting OS, achieving higher C-index values than pathological T stage (training: 0.902 vs. 0.793, p < 0.001; validation: 0.864 vs. 0.757, p = 0.014; external test: 0.841 vs. 0.760, p = 0.047). In addition, MF-DLM-based low-risk group demonstrated significantly longer OS in the training, validation, and external test sets (p < 0.001). In the adjuvant therapy (AT) cohort, high-risk patients had significantly worse prognosis compared with low-risk patients (p < 0.0001). Additionally, high-risk pathological T3/4 patients exhibited no statistically significant OS difference between those who received AT and those who did not (p = 0.18), whereas low-risk pathological T3/4 patients experienced significantly improved OS with AT (p = 0.0059). Besides, the low-risk group had better OS than the high-risk group in neoadjuvant therapy cohort (p = 0.0032). Interpretation: The MF-DLM can reliably predict OS in patients with BCa and provide additional prognostic stratification beyond pathological T and N stages. Furthermore, MF-DLM-based risk groups can identify patients most likely to benefit from perioperative therapy. Funding: The Noncommunicated Chronic Diseases-National Science and Technology Major Project (2024ZD0525700); National Natural Science Foundation of China (82273152, 82503879), Jiangsu Province Hospital (the First Affiliated Hospital of Nanjing Medical University) Clinical Capacity Enhancement Project (JSPH-MA-2022-5), China Postdoctoral Science Foundation funded project (2024M761211), and the Nanjing Postdoctoral Science Foundation funded project (2024BHS210).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".