Comprehensive molecular characterization and analysis of muscle-invasive urothelial carcinomas.
Bibliographic record
Abstract
4500 Background: We reported the integrated molecular analysis of 131 tumors in 2014 (Nature 507:315, 2014) and now report on the entire cohort of 412 tumors from the TCGA project in chemotherapy-naïve, muscle-invasive urothelial bladder cancer. Methods: Following strict clinical and pathologic quality control, tumors were analyzed for DNA copy number variants, somatic mutations (WES), DNA methylation, mRNA, non-coding RNA (lncRNA and miRNA) and (phospho-) protein expression, gene fusions, viral integration, pathway perturbation, clinical correlates, outcomes, and histopathology. Results: There was a high overall somatic mutation rate (8.2/Mb), as previously reported. There were 58 significantly mutated genes (SMGs) (MutSig_2CV), increased from 32 in the original report. We identified 5 mutation signatures including APOBEC-a and b, ERCC2, C > T_CpG, and a single ultra-mutated sample with a functional POLE mutation. APOBEC mutagenesis explained 70% of the mutation burden and was associated with survival (p = 0.0013). High mutation burden and neoantigen load were also associated with improved outcome (p = 0.00014 and 0.00078). The previously identified four mRNA subtypes were predicted on the larger set and also identified a novel poor-survival ‘neuronal’ subtype that nevertheless lacked small cell or neuroendocrine histology. Clustering converged for mRNA, lncRNA and miRNA expression, and for inferred activity of gene sets associated with regulator expression. We identified subsets with differential epithelial-mesenchymal transition scores, carcinoma-in-situ scores, and survival, with implications for distinct therapeutic potential. Conclusions: This integrated analysis of 412 TCGA patient samples validates and extends observations from the first 131 patients and significantly increases our power to detect additional low-frequency aberrations. The results provide unique insights into mechanisms of bladder cancer development, and identify novel subsets of MIBC that may benefit from differential treatment approaches.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".