Abstract A022: Deep Learning Meets Transcriptomics: CNN-1D Powered Interpretation of probe based transcriptome array in Prostate Cancer
Bibliographic record
Abstract
Abstract Precision oncology relies on accurate molecular profiling to guide personalized treatment strategies. The GeneChip™ Human Transcriptome Array 2.0 (HTA 2.0) provides comprehensive whole-transcriptome coverage, enabling simultaneous analysis of gene expression, alternative splicing, and non-coding RNA activity. This high-resolution dataset allows oncologists and researchers to uncover transcriptomic alterations associated with cancer progression, therapy resistance, and disease subtypes. However, the volume and complexity of transcriptomic data require advanced analytical methods for meaningful clinical interpretation. To address this, we have developed a custom AI-powered tool that integrates HTA 2.0 data with a one-dimensional Convolutional Neural Network (CNN-1D) architecture, specifically designed to process sequential gene expression data. The CNN-1D model excels at identifying spatial expression patterns and subtle transcriptomic signatures that correlate with disease phenotypes and therapeutic outcomes. By leveraging this deep learning framework, our tool enables cancer subtype classification, biomarker discovery, and drug response prediction with high accuracy and interpretability. We present a case study focused on prostate cancer, where the tool was trained and validated on HTA 2.0 expression profiles from tumor and normal tissues. The CNN-1D model accurately distinguished between aggressive and indolent prostate cancer subtypes, revealing key splice variants and non-coding RNAs linked to progression and androgen resistance. This integrated AI framework further supports clinical trial design by identifying actionable biomarkers for drug repurposing and determining the viability of immune therapy, especially considering the challenges of targeting solid tumors, which are typically immune "cold". By offering insights into immunogenicity and potential therapeutic targets, the tool helps guide "go/no-go" decisions for utilizing immune therapies in the treatment of solid tumors. Citation Format: Neetu Singh, David Annan, Prashant Bhavar. Deep Learning Meets Transcriptomics: CNN-1D Powered Interpretation of probe based transcriptome array in Prostate Cancer [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A022.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".