Molecular Classification of Epithelial Ovarian Cancer Based on Methylation Profiling: Evidence for Survival Heterogeneity
Bibliographic record
Abstract
Abstract Purpose: Ovarian cancer is a heterogeneous disease that can be divided into multiple subtypes with variable etiology, pathogenesis, and prognosis. We analyzed DNA methylation profiling data to identify biologic subgroups of ovarian cancer and study their relationship with histologic subtypes, copy number variation, RNA expression data, and outcomes. Experimental Design: A total of 162 paraffin-embedded ovarian epithelial tumor tissues, including the five major epithelial ovarian tumor subtypes (high- and low-grade serous, endometrioid, mucinous, and clear cell) and tumors of low malignant potential were selected from two different sources: The Polish Ovarian Cancer study, and the Surveillance, Epidemiology, and End Results Residual Tissue Repository (SEER RTR). Analyses were restricted to Caucasian women. Methylation profiling was conducted using the Illumina 450K methylation array. For 45 tumors array copy number data were available. NanoString gene expression data for 39 genes were available for 61 high-grade serous carcinomas (HGSC). Results: Consensus nonnegative matrix factorization clustering of the 1,000 most variable CpG sites showed four major clusters among all epithelial ovarian cancers. We observed statistically significant differences in survival (log-rank test, P = 9.1 × 10−7) and genomic instability across these clusters. Within HGSC, clustering showed three subgroups with survival differences (log-rank test, P = 0.002). Comparing models with and without methylation subgroups in addition to previously identified gene expression subtypes suggested that the methylation subgroups added significant survival information (P = 0.007). Conclusions: DNA methylation profiling of ovarian cancer identified novel molecular subgroups that had significant survival difference and provided insights into the molecular underpinnings of ovarian cancer. See related commentary by Ishak et al., p. 5729
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".