DNA methylation profile in CpG-depleted regions uncovers a high-risk subtype of early-stage colorectal cancer
Bibliographic record
Abstract
BACKGROUND: The current risk stratification system defined by clinicopathological features does not identify the risk of recurrence in early-stage (stage I-II) colorectal cancer (CRC) with sufficient accuracy. We aimed to investigate whether DNA methylation could serve as a novel biomarker for predicting prognosis in early-stage CRC patients. METHODS: We analyzed the genome-wide methylation status of CpG loci using Infinium MethylationEPIC array run on primary tumor tissues and normal mucosa of early-stage CRC patients to identify potential methylation markers for prognosis. The machine-learning approach was applied to construct a DNA methylation-based prognostic classifier for early-stage CRC (MePEC) using the 4 gene methylation markers FAT3, KAZN, TLE4, and DUSP3. The prognostic value of the classifier was evaluated in 2 independent cohorts (n = 438 and 359, respectively). RESULTS: The comprehensive analysis identified an epigenetic subtype with high risk of recurrence based on a group of CpG loci in the CpG-depleted region. In multivariable analysis, the MePEC classifier was independently and statistically significantly associated with time to recurrence in validation cohort 1 (hazard ratio = 2.35, 95% confidence interval = 1.47 to 3.76, P < .001) and cohort 2 (hazard ratio = 3.20, 95% confidence interval = 1.92 to 5.33, P < .001). All results were further confirmed after each cohort was stratified by clinicopathological variables and molecular subtypes. CONCLUSIONS: We demonstrated the prognostic statistical significance of a DNA methylation profile in the CpG-depleted region, which may serve as a valuable source for tumor biomarkers. MePEC could identify an epigenetic subtype with high risk of recurrence and improve the prognostic accuracy of current clinical variables in early-stage CRC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".