Deep Genomic Signature for early metastasis prediction in prostate cancer
Bibliographic record
Abstract
Abstract For prostate cancer patients, timing and intensity of therapy are adjusted based on their prognosis. Clinical and pathological factors, and recently, gene expression-based signatures have been shown to predict metastatic prostate cancer. Previous studies used labelled datasets, i.e. those with information on the metastasis outcome, to discover gene signatures to predict metastasis. Due to steady progression of prostate cancer, datasets for this cancer have a limited number of labelled samples but more unlabelled samples. In addition to this issue, the high dimensionality of the gene expression data also poses a significant challenge to train a classifier and predict metastasis accurately. In this study, we aim to boost the prediction accuracy by utilizing both labelled and unlabelled datasets together. We propose Deep Genomic Signature (DGS), a method based on Denoising Auto-Encoders (DAEs) and transfer learning. DGS has the following steps: first, we train a DAE on a large unlabelled gene expression dataset to extract the most salient features of its samples. Then, we train another DAE on a small labelled dataset for a similar purpose. Since the labelled dataset is small, we employ a transfer learning approach and use the parameters learned from the first DAE in the second one. This approach enables us to train a large DAE on a small dataset. After training the second DAE, we obtain the list of genes with high weights by applying a standard deviation filter on the transferred and learned weights. Finally, we train an elastic net logistic regression model on the expression of the selected genes to predict metastasis. Because of the elastic net regularization, some of the selected genes have non-zero coefficients in the classifier which we consider as the DGS gene signature for metastasis. We apply DGS to six labelled and one large unlabelled prostate cancer datasets. Results on five validation datasets indicate that DGS outperforms state-of-the-art gene signatures (obtained from only labelled datasets) in terms of prediction accuracy. Survival analyses demonstrate the potential clinical utility of our gene signature that adds novel prognostic information to the well-established clinical factors and the state-of-the-art gene signatures. Finally, pathway analysis reveals that the DGS gene signature captures the hallmarks of prostate cancer metastasis. These results suggest that our method helps to identify a robust gene signature that may improve patient management.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".