Abstract PO-066: Data standardization, integration and meta-analysis of preclinical pharmacogenomics studies for gene expression biomarker discovery
Bibliographic record
Abstract
Abstract The sparsity of predictive biomarkers for drug response remains a major impediment to the clinical success of precision oncology. Over the past decade, massive studies combining in-vitro drug screening with high-throughput molecular profiling of cancer cell lines have been published, with the goal of discovering novel predictive biomarkers for drug response. Unfortunately, inconsistencies in the data, as well as a lack of standards in the field, has effectively siloed the findings of these studies away from the clinical researchers working in precision medicine. In our previous work developing the PharmacoGx/PharmacoDB platform, we have reprocessed and curated the 7 largest pharmacogenomic studies released to date. Together, these studies profiled 760 compounds across ~1600 cell lines, in over 650,000 drug dose response experiments. Here, we present an update of this database, and describe our preliminary findings from a statistical meta-analysis of these data for the discovery of gene expression biomarkers. Our meta-analysis looked for markers consistently associated with drug response across 11 tissue types and 70 compounds found in at least 3 studies in the examined dataset. From these data, we find approximately 4338 consistent associations across studies surviving multiple hypothesis correction, in 1946 genes, across 8 tissues and 34 different drugs. The results show evidence of both drug-mechanism specific markers of response, as well as markers of multi-drug resistance phenotypes. Within lung cancer cell lines, we observe an enrichment of TGF-β signalling related genesets among markers of multi-drug sensitivity and resistance. We discuss our approach to validate these markers using organoid, in vivo and clinical datasets. Finally, we detail tools we are building for researchers to explore our database of putative biomarkers and the data supporting them; focusing on exposing our findings to the clinical precision medicine research community. Ultimately, our goal is to extract a database of high confidence preclinical associations from the consensus of published pharmacogenomic studies. This database would be the first step in translating findings from these massive preclinical studies towards novel clinically actionable biomarkers for use in the pursuit of precision cancer care. Citation Format: Petr Smirnov, Benjamin Haibe-Kains. Data standardization, integration and meta-analysis of preclinical pharmacogenomics studies for gene expression biomarker discovery [abstract]. In: Proceedings of the AACR Virtual Special Conference on Artificial Intelligence, Diagnosis, and Imaging; 2021 Jan 13-14. Philadelphia (PA): AACR; Clin Cancer Res 2021;27(5_Suppl):Abstract nr PO-066.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.133 | 0.262 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.007 | 0.023 |
| Bibliometrics | 0.009 | 0.015 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.006 | 0.003 |
| Open science | 0.005 | 0.005 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".