MOGEDN: small-sample cancer subtype classification with encoder–decoder networks for missing-omics recovery and biomarker discovery
Bibliographic record
Abstract
Effective cancer subtype classification from multi-omics data remains challenging due to incomplete omics data and limited sample sizes. While graph convolutional networks (GCNs) have been used to incorporate inter-sample relationships for enhancing small-sample classification, their performance deteriorates when a certain omics modality is entirely missing. Here, we propose MOGEDN, a novel framework for cancer subtype classification using multi-omics encoder-decoder networks designed to reconstruct the latent features of missing omics data. The reconstructed features are integrated with available omics features to enable robust prediction under small-sample and missing-omics settings. We develop a step-wise algorithm to pretrain our model with diverse cancer types then to finetune for a specific cancer type while incorporating inter-sample and cross-omics dependencies. Evaluated on TCGA cancer datasets including subtypes with fewer than 50 samples, MOGEDEN consistently outperforms state-of-the-art baselines in accuracy and F1 scores. Moreover, MOGEDN's feature analysis provides two complementary biomarker sets: biomarkers shared across diverse cancer types in the pretraining phase; and biomarkers for a specific cancer type in the finetuning phase, facilitating model interpretability, and biological findings. These results highlight decoder-based imputation as a powerful approach to enhance multi-omics learning, delivering accurate classification, robust few-shot performance, and multi-scale biomarker discovery in incomplete multi-omics cohorts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".