Abstract A011: Interpretable machine learning for discovery, evaluation and clinical translation of context-specific dependencies
Bibliographic record
Abstract
Abstract Functional genomics and compound screens across cancer cell line models have identified thousands of selective vulnerabilities, either gene dependencies or drug responses. Linking these vulnerabilities to genetic, epigenetic and molecular states that drive tumorigenesis is essential to develop them as therapeutic targets. While simple relationships are easily captured by binary or linear associations, such as a driver gene and its mutation or expression, complex or heterogenous relationships are more difficult to characterize. Machine learning approaches are commonly employed to associate vulnerabilities with molecular contexts; however, off-the-shelf models lack the comprehensive explainability required to effectively capture complex biological context. To address this shortcoming, we developed Oncoforest, a machine learning toolkit designed for biologically interpretable prediction of viability readouts and association with molecular context. Oncoforest implements forest-based model variants designed for explainability from large feature-sets, and builds biologically-informed attribution networks from model Shapley values. Using these models and networks, we successfully recapitulate the known molecular context of driver genes and propose hundreds of novel context-specific vulnerabilities. We further categorize these into types based on the context relationship, for example, self gain-of-function, paralog loss-of-function, or other synthetic or collateral lethality. In order to consider the therapeutic potential of gene expression-based molecular context, it is important to be able to evaluate their translation to patient tumors. To enable this, we have generated a harmonized transcriptomic map consisting of 1019 cancer cell lines, 14604 tumors (from TCGA and clinical trials), and 9635 normals. Applying the models to predict viability readouts within the tumor and normal bulk-tissue map, Oncoforest estimates cancer-vulnerability and normal-tissue sensitivity of selective gene dependencies and drug responses. For greater granularity, learned molecular signatures were applied to score 3146 single-cell types from a single-cell atlas. We demonstrate that these predictions are able to identify known tissue and cell-type sensitivities from therapeutic intervention, as well as to stratify patient outcomes, underscoring the clinical relevance of such approaches. Taken together, Oncoforest represents a set of machine learning tools for enhanced interpretability, suitable for discovery and assessment of cancer vulnerabilities. Citation Format: James Mondo, Michal Kabza, Julien Tremblay, Barzin Nabet, Marc Hafner, Xiaosai Yao, Timothy Sterne-Weiler. Interpretable machine learning for discovery, evaluation and clinical translation of context-specific dependencies [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr A011.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.019 | 0.005 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".