MétaCan
Menu
Back to cohort
Record W2576982189 · doi:10.1016/j.ebiom.2017.02.022

Mining Human Prostate Cancer Datasets: The “camcAPP” Shiny App

2017· article· en· W2576982189 on OpenAlexaff
Mark Dunning, Sarah L. Vowler, Emilie Lalonde, Helen Ross‐Adams, Paul C. Boutros, Ian G. Mills, Andy G. Lynch, Alastair Lamb

Bibliographic record

VenueEBioMedicine · 2017
Typearticle
Languageen
FieldMedicine
TopicProstate Cancer Treatment and Research
Canadian institutionsOntario Institute for Cancer ResearchUniversity of Toronto
FundersAcademy of Medical SciencesNational Institute for Health and Care ResearchCancer Research UK
KeywordsProstate cancerComputer scienceProstatectomyAnnotationRelevance (law)CancerInformation retrievalBioinformaticsMedicineArtificial intelligenceBiologyInternal medicine

Abstract

fetched live from OpenAlex

Obtaining access to robust, well-annotated human genomic datasets is an important step in demonstrating the relevance of experimental findings and, often, in generating the hypotheses that led to those experiments being conducted in the first place. We recently published data from the CamCaP Study Group which comprised two cohorts of men with prostate cancer who had undergone prostatectomy in Cambridge, UK and Stockholm, Sweden (Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google Scholar). We considered how we might best share our output with those who wish to interrogate the data with their own ideas, gene lists and clinical questions. We recognised that finding, down-loading, pre-processing and assimilating any such dataset into a usable format is daunting and may put off many researchers. We also felt that interrogation tools generated to date (e.g. cBioPortal) lack functionality as they either cover too many organ types, or are limited in the extent, precision and tumour-site specificity of their clinical annotation. We therefore determined to produce an accessible web-based platform that would permit straightforward interrogation of these datasets with individual gene identifiers or gene sets. Furthermore, we decided to include additional ‘publicly-accessible’ human prostate cancer sets in order to increase the number of samples available and provide a degree of validation of any observations made across independent cohorts. We included a number of prominent publicly available sets with both gene expression and copy number data leading to a cohort of almost 500 men (Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google Scholar, Taylor et al., 2010Taylor B.S. Schultz N. Hieronymus H. Gopalan A. Xiao Y. Carver B.S. et al.Integrative genomic profiling of human prostate cancer.Cancer Cell. 2010; 18: 11-22Summary Full Text Full Text PDF PubMed Scopus (2740) Google Scholar, Grasso et al., 2012Grasso C.S. Wu Y.M. Robinson D.R. Cao X. Dhanasekaran S.M. Khan A.P. et al.The mutational landscape of lethal castration-resistant prostate cancer.Nature. 2012; 487: 239-243Crossref PubMed Scopus (1800) Google Scholar). We also included a small landmark series of expression data (Varambally et al., 2005Varambally S. Yu J. Laxman B. Rhodes D.R. Mehra R. Tomlins S.A. et al.Integrative genomic and proteomic analysis of prostate cancer reveals signatures of metastatic progression.Cancer Cell. 2005; 8: 393-406Summary Full Text Full Text PDF PubMed Scopus (638) Google Scholar). These studies are summarised in Table 1. We plan to include additional studies in the app as well-annotated datasets become publicly available.Table 1Summary of studies included in the camcAPP at initial release. Primary Tumours = tissue taken from radical prostatectomy specimens in men with confirmed organ-confined disease. Advanced Tumours = tissue from channel transurethral resection of the prostate (chTURP) or prostatectomy in men with metastatic disease.DatasetPaperPlatform: gene expressionPlatform: copy numberPrimary tumoursAdvanced tumoursClinical covariatesMichigan 2005Varambally et al., 2005Varambally S. Yu J. Laxman B. Rhodes D.R. Mehra R. Tomlins S.A. et al.Integrative genomic and proteomic analysis of prostate cancer reveals signatures of metastatic progression.Cancer Cell. 2005; 8: 393-406Summary Full Text Full Text PDF PubMed Scopus (638) Google ScholarGSE3325Affymetrix U133 2.0N/A76Sample GroupMSKCC 2010Taylor et al., 2010Taylor B.S. Schultz N. Hieronymus H. Gopalan A. Xiao Y. Carver B.S. et al.Integrative genomic profiling of human prostate cancer.Cancer Cell. 2010; 18: 11-22Summary Full Text Full Text PDF PubMed Scopus (2740) Google ScholarGSE21032Affymetrix Human 1.0 STAgilent 244k10919Gleason, Copy Number ClusterMichigan 2012Grasso et al., 2012Grasso C.S. Wu Y.M. Robinson D.R. Cao X. Dhanasekaran S.M. Khan A.P. et al.The mutational landscape of lethal castration-resistant prostate cancer.Nature. 2012; 487: 239-243Crossref PubMed Scopus (1800) Google ScholarGSE35988Agilent Whole Human 44kAgilent 105k/244k5932Sample GroupStockholm 2015Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google ScholarGSE70770Illumina HT12Affymetrix SNP 6.0101N/AiCluster, Sample GroupCambridge 2015Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google ScholarGSE70770Illumina HT12Illumina Omni 2.512519iCluster, Gleason, Sample Group Open table in a new tab An important finding in our recent study was that prostate cancer could be divided into five distinct molecular subgroups based on stratification with a small number of copy number features which were also associated with RNA-expression change. These groups had different clinical outcomes. We wanted the app to allow researchers to determine the mean RNA-expression level or copy number status of a single gene or gene-set in prostates from men divided either according to clinical categories (Gleason score, biochemical relapse status or tumour type) or according to molecular subgroups. These subgroups could either be pre-defined molecular groups published in the relevant papers, or de novo subgroups generated by hierarchical clustering based on an uploaded geneset. We searched for other web-tools that are already available for this purpose. Although no such site exists for assessment of subgroup patterns or combined expression and copy number profiles, the Memorial Sloane Kettering Cancer Centre (MSKCC) and Michigan data (Table 1) can be analysed as part of cBioPortal (cBioPortal for Cancer Genomics, n.dcBioPortal for Cancer Genomics. Memorial Sloane Kettering Cancer Centre, www.cbioportal.org, (Accessed: 23/03/2016).Google Scholar) along with the recently published prostate TCGA dataset (Robinson et al., 2015Robinson D. Van Allen E.M. Wu Y.M. Schultz N. Lonigro R.J. Mosquera J.M. et al.Integrative clinical genomics of advanced prostate cancer.Cell. 2015; 161: 1215-1228Summary Full Text Full Text PDF PubMed Scopus (2011) Google Scholar). Here we introduce the camcAPP (http://bioinformatics.cruk.cam.ac.uk/apps/camcAPP/); a bespoke web interface to multiple prostate cancer genomics datasets. The interface was created with Shiny (https://www.rstudio.com/products/shiny/), and allows the non-specialist Bioinformatician to create publication-ready figures and tables through an intuitive interface to the underlying computer code. After selecting a dataset of interest, and uploading a list of genes, the following analyses can be performed:1)Boxplots and analysis of variance for expression of genes of interest grouped by clinical group, sample type, Gleason grade of copy-number cluster (N.B. not all covariates available for all data sets) (see Supplementary Fig. 2).2)A recursive partitioning-based survival analysis and Kaplan-Meier plots on a gene-by-gene basis (Supplementary Fig. 3).3)Pairwise-correlations of gene expression across whole studies and within clinical subgroups.4)Clustering and heatmaps of gene expression data, with options to interrogate associations between clinical covariates and newly-derived clusters (Supplementary Fig. 4).5)Tabulating the number of copy-number amplifications and deletions observed across whole studies or within a particular clinical covariate, and making a heatmap of copy-number calls (Supplementary Fig. 5) (Lalonde et al., 2014Lalonde E. Ishkanian A.S. Sykes J. Fraser M. Ross-Adams H. Erho N. et al.Tumour genomic and microenvironmental heterogeneity for integrated prediction of 5-year biochemical recurrence of prostate cancer: a retrospective cohort study.Lancet Oncol. 2014; 15: 1521-1532Summary Full Text Full Text PDF PubMed Scopus (236) Google Scholar). One of the challenges in constructing such a tool is delivering an output format that is readily transferable to slides for presentations or panels of a figure for publication. We recognise that this is, in part, a matter of axis typesetting and plot configuration but also of delivering an output file which permits further adjustment of the figure in, for example, Adobe Illustrator™. To this end, all plots can be exported as PDF or PNG files with configurable dimensions. Furthermore, for those that are well-versed in R, the code to produce a particular plot can be downloaded and modified as required. A further challenge that we seek to address with this interface is merging datasets for combined analysis. We hope to offer this option in due course, as we include further datasets that include samples analysed on compatible platforms. Strategies to address the Big Data problem have focussed on making the ever-increasing volume of genomic data accessible to scientists and on opening up the possibility of engaging non-specialists (Keener, 2015Keener Amanda B. The Scientist. July 8, 2015; (Accessed: 23/3/16)http://www.the-scientist.com/?articles.view/articleNo/43483/title/Big-Data-Problem/Google Scholar). This approach embodies a responsible attitude to science both in terms of patient input and financial resource and we believe that tools such as this are an important step to maximising the value of these landmark studies. We take pleasure in making this platform available to the prostate cancer community by means of this ‘In Focus’ article in EBioMedicine, a journal that we believe champions this responsible approach to genomic data both in cancer genomics (Kerns et al., 2016Kerns S.L. Dorling L. Fachal L. Bentzen S. Pharoah P.D. Barnes D.R. et al.Meta-analysis of genome wide association studies identifies genetic markers of late toxicity following radiotherapy for prostate cancer: radiogenomics consortium.EBioMedicine. 2016; 10: 150-163https://doi.org/10.1016/j.ebiom.2016.07.022Summary Full Text Full Text PDF PubMed Scopus (56) Google Scholar) and further afield (Taudien et al., 2016Taudien S. Lausser L. Giamarellos-Bourboulis E.J. Sponholz C. Schöneweck F. Felder M. et al.Genetic factors of the disease course after sepsis: rare deleterious variants are predictive.EBioMedicine. 2016; 12: 227-238https://doi.org/10.1016/j.ebiom.2016.08.037Summary Full Text Full Text PDF PubMed Scopus (27) Google Scholar). Core CRUK funding: MD, AGL, ADL.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.123
Threshold uncertainty score0.802

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0010.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.056
GPT teacher head0.410
Teacher spread0.355 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations37
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueEBioMedicineSame topicProstate Cancer Treatment and ResearchFrench-language works237,207