Identification of CYP2B6 as a Novel Biomarker of HRD in Colon Adenocarcinoma through WGCNA and Machine Learning
Bibliographic record
Abstract
The potential role of homologous recombination deficiency (HRD) in the diagnosis and treatment of colon adenocarcinoma (COAD) remains incompletely explored. Differential gene expression analysis was conducted using Limma to identify genes with altered expression levels. Key genes associated with HRD were identified through the integration of WGCNA and machine learning techniques. For the unsupervised grouping of samples, ConsensusClusterPlus was applied. To quantify gene expression and protein abundance in clinical tissues and cell lines, RT-qPCR and Western Blotting (WB) assays were performed, respectively. The "pRRophetic" package was employed to predict drug sensitivity profiles. Molecular docking simulations and optimal pose presentations were conducted by using CB-Dock2. Our comprehensive analysis of multiple COAD data sets, leveraging WGCNA and machine learning, unveiled five novel, previously unreported biomarkers of HRD: TNFRSF11A, SERPINA1, SPINK4, REG4, and CYP2B6. We devised an innovative HRD-linked molecular classification system and a predictive nomogram that accurately forecasts patient outcomes. Experimental validation substantiated the upregulation of CYP2B6 in COAD, enhancing proliferation and migration capabilities, and demonstrated a robust positive association with established HRD indicators RAD51 and γH2AX. Notably, CYP2B6 emerged as a promising predictor of PARP inhibitor (PARPi) sensitivity, offering potential therapeutic implications. In conclusion, our study, harnessing machine learning and experimental validation, has uncovered novel biomarkers of HRD and PARPi sensitivity, shedding light on potential avenues for tailored clinical treatment strategies in COAD, thereby advancing personalized medicine.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".