Abstract 4930: Identification of CBX7 and PCDHB18 as novel prognostic biomarkers of cervical cancer: RNA-sequencing and machine learning analysis
Bibliographic record
Abstract
Abstract Background- Cervical cancer, ranking fourth in prevalence among women worldwide, is highly preventable when diagnosed early. The advent of innovative technologies including bioinformatics and machine learning, is revolutionizing the discovery and development of novel biomarkers for cancer. Our study uniquely integrates artificial neural networks with RNA-sequencing data of cervical cancer patients. This combination enhances the accuracy and reliability of biomarker predictions, providing a better understanding of the potential clinical utility of the identified biomarkers. Methods- RNA-sequencing data and clinicopathological details of 304 Cervical Squamous Cell Carcinoma and Endocervical Adenocarcinoma (CESC) were obtained from the GDAC database (https://gdac.broadinstitute.org/). Differentially expressed genes (DEGs) were identified (P<0.05, |log2 fold change (FC)| > 1.5, false discovery rate (FDR) < 0.05). Pathway enrichment analysis was performed and protein-protein interactions of DEGs were constructed using the STRING database. Prognostic biomarkers were identified using Kaplan-Meier and Cox proportional hazard methods adjusted for potential confounders such as age, disease stage, and comorbidities (HR<1, p<0.05). Deep learning algorithms were applied for predictive marker identification, utilizing Weight by Correlation feature selection. The model used AUC (Area Under the Curve), accuracy, MSE (Mean Squared Error), and R2 (R-squared) as evaluation metrics in a 70-30 training-test split. A combined ROC curve assessed diagnostic biomarkers, validated externally with a GEO dataset on cervical cancer patients. Results- In our study, 4153 DEGs were identified. Pathway analysis revealed that key dysregulated genes play a pivotal role in extracellular matrix organization. The survival analysis identified that six upregulated genes (CT62, SLC7A5P1, C10orf110, FDPSL2A, TNFSF15, and TTLL13) and seven downregulated genes (BAI3, CBX7, DKFZp566F0947, PCDHB18, GRAPL, PCDHB19P, and C4orf38) decreased the overall survival in patients. The machine learning model demonstrated high predictive accuracy (AUC=1, accuracy=99.02%, R2=0.99), identifying twenty genes with a positive correlation to cervical cancer risk. Notably, CBX7 emerged as a prognostic and diagnostic biomarker (AUC=0.99, sensitivity=0.93, specificity=1.00). Conclusion- Our study uncovers the prognostic significance of two novel genes in cervical cancer: CBX7, a vital regulator of tumor suppression, and PCDHB18, a member of the protocadherin beta gene cluster functioning as a cell adhesion molecule. Downregulation of these genes is associated with decreased overall survival. Further functional analyses and validation of these candidate biomarkers are crucial to fully assess their potential clinical value in cervical cancer. Citation Format: Ghazaleh Pourali, Mohsen Zeinali, Mahshid Arastonejad, Nima Khalili-Tanha, Elham Nazari, Ghazaleh Khalili-Tanha, Adetunji T. Toriola. Identification of CBX7 and PCDHB18 as novel prognostic biomarkers of cervical cancer: RNA-sequencing and machine learning analysis [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 4930.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".