Abstract A002: Computational identification for missense mutations of bladder cancer candidate genes using TCGA datasets
Bibliographic record
Abstract
Abstract Bladder cancer is the fourth most common cancer in men. According to the American Cancer Society, there were roughly over 80,000 cases in the United States in 2023, and nine out of ten people with this cancer are over the age of 55. The Cancer Genome Atlas (TCGA) project has made it possible for researchers to study mutations and other types of genetic signatures for multiple cancers including Bladder Urothelial Carcinoma (BLCA). A missense mutation is where a single nucleotide change in the DNA sequence results in the substitution of one amino acid for another in the protein. This could lead to a protein with an altered function. The impact of a missense mutation could be benign or harmful. In our previous study, 54 significant clusters for 15 cancer types were identified in The Cancer Genome Atlas (TCGA) network. In this study, we further investigated genes in two significant clusters reported for BLCA. We initially aim to identify genes that have missense mutations that have functional impact. To achieve this goal, we plugged the two clusters of genes (199) into the cBioPortal website and identified eight genes that had missense mutations or met three criteria (Firstly, numbers of mutations in sample must be greater than or equal to 5; secondly, functional impact must meet the following both: SIFT impact is deleterious and Polyphen-2 impact is possibly damaging or probably damaging): AKT3, PARP1, PRKN, APC, GNA11, KMT2A, NEGR1, and FYN. Further analysis showed what specific amino acid had been substituted at what specific position on the sequence. For example, for gene AKT3, at position 286, amino acid K had been substituted into amino acid N. Similarly, for gene PARP1, at position 913, amino acid G had been substituted into amino acid A. We then identified every amino acid missense substitution for all eight genes. By examining the mutation pattern, we hope to identify key genes as biomarkers so that therapeutic efficacy can be optimized. Citation Format: Kai He, Lucas Wang, Yongsheng Bai. Computational identification for missense mutations of bladder cancer candidate genes using TCGA datasets. [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Optimizing Therapeutic Efficacy and Tolerability through Cancer Chemistry; 2024 Dec 9-11; Toronto, Ontario, Canada. Philadelphia (PA): AACR; Mol Cancer Ther 2024;23(12_Suppl):Abstract nr A002
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.006 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".