MétaCan
Menu
← Back to cohort
Record W2886615727 · doi:10.1158/1538-7445.am2018-1418

Abstract 1418: Identification of recurrent regulatory mutations in breast cancer

2018· article· en· W2886615727 on OpenAlexaff
Kelsy C. Cotto, Arpad Danos, Robert Lesurf, Morag Park, Malachi Griffith, Obi L. Griffith

Bibliographic record

VenueCancer Research · 2018
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsConcordia UniversityMcGill UniversityOntario Institute for Cancer Research
Fundersnot available
KeywordsBreast cancerBiologyTranscriptomeExomeGeneGeneticsComputational biologyCancerExome sequencingPTENPromoterPhenotypeCancer researchGene expression

Abstract

fetched live from OpenAlex

Abstract Since the identification of recurrent TERT promoter mutations in melanoma resulting in increased TERT expression, there has been increased interest in identifying recurrent regulatory non-coding mutations (Horn et al. 2013, Huang et al. 2013). Several studies have attempted pan-cancer analyses in order to identify these types of mutations, but often the results suffer from low coverage of regulatory regions or do not extend to breast cancer. While some breast cancer specific studies have identified some significantly mutated promoters and lncRNAs, they have often failed to incorporate transcriptome data to assess the impact and relevance of mutations on the expression of genes within tumors (Nik-Zainal et al. 2016). In order to address this, we assembled and generated a data set consisting of 458 breast cancer cases with matched tumor/normal pairs. This cohort consists of 22.4% luminal A, 19% luminal B, 16.4% HER2-enriched, 21% basal-like, 0.8% normal-like, and 20.4% unknown with regards to molecular subtype. This is important due to different breast cancer subtypes having dissimilar phenotypes and varying rates of gene coding mutations. This data set has a mix of whole genome, exome, transcriptome, and custom capture sequencing. We designed a custom capture reagent that covers regions assembled from regulatory databases, 5' untranslated regions, 500 bases upstream and downstream of transcription start sites, and 50,000 bases upstream and downstream of 178 genes that have been implicated as being important in breast cancer (Lesurf et al. 2016). While this custom capture region is similar in size to an exome, it has advantages over whole genome and exome sequencing, particularly with respect to coverage in GC-rich promoter regions. With these data, we predict that we will be able to identify novel, regulatory coding and non-coding drivers of breast cancer that would not be discovered without integrated analysis of the DNA- and RNA-seq data for each tumor. Instrument data were processed using the McDonnell Genome Institute somatic variant calling pipeline that includes 5 SNV callers and 3 indel callers. We then used these steps to filter variants: min. 20x coverage in both the tumor and normal sample, min. 2.5% tumor variant allele frequency, min. 3 variant supporting reads in the tumor sample, max. 10% variant allele frequency in the normal sample. We also filtered against gnomAD and a panel of normals. Rheinbay et al. 2017 identified recurrently mutated promoter regions for nine genes: TBC1D12, ZNF143, ALDOA, NEAT1, RMRP, CITED2, FOXA1, CTNNB1, LEPROTL1. Our preliminary analysis has revealed that we also see mutations within these regions. We plan to present on the significance of mutations within these previously seen regions, based on recurrence and transcriptome changes, as well as novel recurrent regulatory regions that our analysis reveals, particularly with respect to molecular subtype. Citation Format: Kelsy C. Cotto, Arpad Danos, Robert Lesurf, Morag Park, Malachi Griffith, Obi L. Griffith. Identification of recurrent regulatory mutations in breast cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1418.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.002
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.004
Threshold uncertainty score0.014

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0040.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.041
GPT teacher head0.396
Teacher spread0.354 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2018
Admission routes1
Has abstractyes

Explore more

Same venueCancer Research→Same topicCancer Genomics and Diagnostics→French-language works237,207→