MétaCan
Menu
Back to cohort
Record W2740244858 · doi:10.1158/1538-7445.am2017-385

Abstract 385: Network-driven discovery of cancer drivers and pathways using 2,500 whole cancer genomes

2017· article· en· W2740244858 on OpenAlexaff
Jüri Reimand

Bibliographic record

VenueCancer Research · 2017
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicRNA modifications and cancer
Canadian institutionsOntario Institute for Cancer Research
Fundersnot available
KeywordsBiologyGeneticsIndelGeneGenomeComputational biologyEnhancerCancerCoding regionMutationUntranslated regionTranscription factorSingle-nucleotide polymorphismMessenger RNA

Abstract

fetched live from OpenAlex

Abstract Cancer driver genes exhibit unexpectedly high mutation rates in large cancer genomic datasets. We hypothesize that driver mutations specifically alter molecular interaction networks by disrupting “active sites” - interaction interfaces in proteins and DNA. We present ActiveDriverWGS, a novel computational method to discover cancer drivers in whole-genome sequencing (WGS) data. ActiveDriverWGS finds genome regions that are significantly enriched in somatic single nucleotide variants (SNVs) and indels and ascertains whether these associate to known active sites. Analysis of active sites allows us to predict the mechanisms of mutation on three layers of the central dogma: regulatory DNA with transcription factor (TF) binding sites (TFBS), mRNA with microRNA binding sites in untranslated regions (UTRs), and post-translational modification (PTM) sites in proteins. To discover cancer driver genes and pathways, we analysed the WGS dataset of >2,500 samples from the International Cancer Genome Consortium (ICGC) Pan-Cancer Analysis Working Group (PCAWG). We found 61 protein-coding candidates with 34 known drivers (P=10-40), validating the high accuracy of our method. 40 genes have significant mutations of PTM sites, suggesting that rewiring of PTM signalling networks is a common oncogenic mechanism. For example, the BRAF V600E SNV flanks two phosphorylation sites and one ubiquitination site (FDR P=10-44), a novel interpretation and potential avenue for precision therapies targeting the kinase and ubiquitin network of BRAF. In the non-coding genome, we detected known lncRNAs (NEAT1, MALAT1), promoters (TERT, WDR74) and novel candidates with mutation enrichment. For example, an enhancer on chr6 has a mutation hotspot in 33 patients (FDR P=10-19), with 20 SNVs affecting binding motifs of cancer-associated TFs FOXO3, SOX2, HMGA2 (FDR P=10-10). Thus our method discovers non-coding drivers and their candidate mechanisms in a single analysis. Our ActiveDriverPW method extends coding and non-coding mutations to biological pathways. We found >600 mutation-enriched pathways in the PCAWG pan-cancer dataset. Of these ~200 are also significant when only non-coding mutations are analysed, showing that the non-coding genome includes previously unstudied mutations in pathways. The DNA double-strand break response pathway (FDR p=10-10) includes non-coding SNVs in ~20 histones and chromatin modifiers, such as the demethylase KDM4B with 46 SNVs in its promoter and enhancers. ActiveDriverPW maps mutations of the long tail that affect genes in hallmark cancer processes yet remain undiscovered in gene-focused analyses. Our methods accurately capture known drivers in the ICGC-PCAWG dataset and suggest specific mechanistic details. Our benchmarks also emphasize the robust performance of our methods. ActiveDriverWGS and ActiveDriverPW are valuable additions to the toolbox for cancer genome analysis. Citation Format: Jüri Reimand. Network-driven discovery of cancer drivers and pathways using 2,500 whole cancer genomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 385. doi:10.1158/1538-7445.AM2017-385

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.004
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.009
Threshold uncertainty score0.019

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.004
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0030.002
Science and technology studies0.0010.000
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0050.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.104
GPT teacher head0.407
Teacher spread0.303 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueCancer ResearchSame topicRNA modifications and cancerFrench-language works237,207