Abstract A001: Harnessing machine learning for the virtual screening of natural compounds as both EGFR and HER2 inhibitors in colorectal Cancer: A novel therapeutic approach
Bibliographic record
Abstract
Abstract Colorectal cancer (CRC) treatment has progressed with the introduction of targeted drugs, particularly those that inhibit the epidermal growth factor receptor (EGFR) and human epidermal growth factor receptor 2 (HER2). EGRF and HER2 are validated targets in colon cancer therapy which are overexpressed in up to 85% of CRC cases. Present studies have focused on single blockage of EGFR or HER2 to combat colon cancer. However, monotherapies that target either EGFR or HER2 receptor frequently have low efficacy due mutations in downstream effectors such as Kirsten Rat Sarcoma 2 Viral Oncogene Homolog (KRAS) and the activation of compensatory signaling pathways that support tumour survival and proliferation. Hence, the discovery and development of a therapy with the capability to combat CRC by simultaneously inhibiting both EGFR and HER2, remains an avenue for further exploration. This study introduces a novel machine learning (ML)-based stacking ensemble framework for rapidly and accurately identifying dual EGFR and HER2 inhibitors using SMILES notation. A benchmark dataset comprising active and inactive compounds against EGFR and HER2 was curated from the ChEMBL database. Based on this dataset, baseline models were developed and optimized using a collection of comprehensive set of well-known molecular descriptors and ML algorithms. These models generated probabilistic features integrated via logistic regression model as a final estimate to construct the final stacking ensemble model. Experimental results indicated that the stacking framework achieved remarkable predictive performance, with accuracy and Matthews correlation coefficient (MCC) values of 0.995 and 0.988 on the training dataset and 100% accuracy and MCC on independent test data. Additionally, molecular docking studies were conducted on the top three compounds predicted by the stacking model and four FDA-approved cancer drugs. Among these, LTS0018035 demonstrated the highest binding affinity against HER2 (PDB ID: 7MN5), with a binding energy of –11.2 kcal/mol and an inhibition constant of 0.00626 μM, outperforming Tucatinib, a standard CRC treatment. This study underscores the potential of ML-driven approaches in accelerating the discovery of dual-target inhibitors for CRC therapy. Citation Format: Deli-Bright Nii Tettey. Oku, Ramakwala Christinah. Chokwe. Harnessing machine learning for the virtual screening of natural compounds as both EGFR and HER2 inhibitors in colorectal Cancer: A novel therapeutic approach [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A001.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".