EPIG-13ROBUST MGMT METHYLATION DETECTION USING 450k ARRAY IN CIMP-NEGATIVE GBM
Bibliographic record
Abstract
Robust and accurate testing of MGMT methylation as a predictive biomarker is of vital importance to predict response to temozolomide in GBM patients. Although various approaches were published previously for DNA methylation based cancer classification, there is a need for improving accuracy, reproducibility and clarity. To achieve this, we set guidelines for reproducibility and created a classification framework for predicting MGMT methylation efficiently using predictive modeling and validation. We have utilized Illumina 450k genome-wide methylation signatures to identify CpGs whose methylation status correlates with MGMT status. We reflected realistic clinical aspect by separating training and validation dataset to avoid biased feature selection, and performed cross validation (2-level 8-fold). Feature extraction was performed by ANOVA to identify biomarkers that reflect the methylation status of MGMT gene for GBM classification. Samples that showed CIMP-positive profiles by 450k were excluded, given the confounding of CIMP status with MGMT methylation. We examined 450k data from 191 samples of GBM which were tested on the Illumina 450k array at 2 centers and compared array data with MGMT methylation status determined by methylation specific PCR (MSP) assay. We selected 5 probes based on a random forest model, where 2 of the 5 probes are in common with the previously reported MGMT-STP27. The RF based model produced 93.5% concordance with MSP, as compared to STP-27, which showed 79% concordance. The discordant samples will be re-assessed with MSP assay to compare the accuracy of MSP with the 450K based approach. While further validation is in progress, this robust framework can efficiently identify additional methylation features correlated with MGMT methylation status and, data from the 450k array can be used to detect MGMT methylation status.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".