Mining Human Prostate Cancer Datasets: The “camcAPP” Shiny App
Notice bibliographique
Résumé
Obtaining access to robust, well-annotated human genomic datasets is an important step in demonstrating the relevance of experimental findings and, often, in generating the hypotheses that led to those experiments being conducted in the first place. We recently published data from the CamCaP Study Group which comprised two cohorts of men with prostate cancer who had undergone prostatectomy in Cambridge, UK and Stockholm, Sweden (Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google Scholar). We considered how we might best share our output with those who wish to interrogate the data with their own ideas, gene lists and clinical questions. We recognised that finding, down-loading, pre-processing and assimilating any such dataset into a usable format is daunting and may put off many researchers. We also felt that interrogation tools generated to date (e.g. cBioPortal) lack functionality as they either cover too many organ types, or are limited in the extent, precision and tumour-site specificity of their clinical annotation. We therefore determined to produce an accessible web-based platform that would permit straightforward interrogation of these datasets with individual gene identifiers or gene sets. Furthermore, we decided to include additional ‘publicly-accessible’ human prostate cancer sets in order to increase the number of samples available and provide a degree of validation of any observations made across independent cohorts. We included a number of prominent publicly available sets with both gene expression and copy number data leading to a cohort of almost 500 men (Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google Scholar, Taylor et al., 2010Taylor B.S. Schultz N. Hieronymus H. Gopalan A. Xiao Y. Carver B.S. et al.Integrative genomic profiling of human prostate cancer.Cancer Cell. 2010; 18: 11-22Summary Full Text Full Text PDF PubMed Scopus (2740) Google Scholar, Grasso et al., 2012Grasso C.S. Wu Y.M. Robinson D.R. Cao X. Dhanasekaran S.M. Khan A.P. et al.The mutational landscape of lethal castration-resistant prostate cancer.Nature. 2012; 487: 239-243Crossref PubMed Scopus (1800) Google Scholar). We also included a small landmark series of expression data (Varambally et al., 2005Varambally S. Yu J. Laxman B. Rhodes D.R. Mehra R. Tomlins S.A. et al.Integrative genomic and proteomic analysis of prostate cancer reveals signatures of metastatic progression.Cancer Cell. 2005; 8: 393-406Summary Full Text Full Text PDF PubMed Scopus (638) Google Scholar). These studies are summarised in Table 1. We plan to include additional studies in the app as well-annotated datasets become publicly available.Table 1Summary of studies included in the camcAPP at initial release. Primary Tumours = tissue taken from radical prostatectomy specimens in men with confirmed organ-confined disease. Advanced Tumours = tissue from channel transurethral resection of the prostate (chTURP) or prostatectomy in men with metastatic disease.DatasetPaperPlatform: gene expressionPlatform: copy numberPrimary tumoursAdvanced tumoursClinical covariatesMichigan 2005Varambally et al., 2005Varambally S. Yu J. Laxman B. Rhodes D.R. Mehra R. Tomlins S.A. et al.Integrative genomic and proteomic analysis of prostate cancer reveals signatures of metastatic progression.Cancer Cell. 2005; 8: 393-406Summary Full Text Full Text PDF PubMed Scopus (638) Google ScholarGSE3325Affymetrix U133 2.0N/A76Sample GroupMSKCC 2010Taylor et al., 2010Taylor B.S. Schultz N. Hieronymus H. Gopalan A. Xiao Y. Carver B.S. et al.Integrative genomic profiling of human prostate cancer.Cancer Cell. 2010; 18: 11-22Summary Full Text Full Text PDF PubMed Scopus (2740) Google ScholarGSE21032Affymetrix Human 1.0 STAgilent 244k10919Gleason, Copy Number ClusterMichigan 2012Grasso et al., 2012Grasso C.S. Wu Y.M. Robinson D.R. Cao X. Dhanasekaran S.M. Khan A.P. et al.The mutational landscape of lethal castration-resistant prostate cancer.Nature. 2012; 487: 239-243Crossref PubMed Scopus (1800) Google ScholarGSE35988Agilent Whole Human 44kAgilent 105k/244k5932Sample GroupStockholm 2015Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google ScholarGSE70770Illumina HT12Affymetrix SNP 6.0101N/AiCluster, Sample GroupCambridge 2015Ross-Adams et al., 2015Ross-Adams H. Lamb A.D. Dunning M.J. Halim S. Lindberg J. Massie C.M. et al.Integration of copy number and transcriptomics provides risk stratification in prostate cancer: a discovery and validation cohort study.EBioMedicine. 2015; 2: 1133-1144Summary Full Text Full Text PDF PubMed Google ScholarGSE70770Illumina HT12Illumina Omni 2.512519iCluster, Gleason, Sample Group Open table in a new tab An important finding in our recent study was that prostate cancer could be divided into five distinct molecular subgroups based on stratification with a small number of copy number features which were also associated with RNA-expression change. These groups had different clinical outcomes. We wanted the app to allow researchers to determine the mean RNA-expression level or copy number status of a single gene or gene-set in prostates from men divided either according to clinical categories (Gleason score, biochemical relapse status or tumour type) or according to molecular subgroups. These subgroups could either be pre-defined molecular groups published in the relevant papers, or de novo subgroups generated by hierarchical clustering based on an uploaded geneset. We searched for other web-tools that are already available for this purpose. Although no such site exists for assessment of subgroup patterns or combined expression and copy number profiles, the Memorial Sloane Kettering Cancer Centre (MSKCC) and Michigan data (Table 1) can be analysed as part of cBioPortal (cBioPortal for Cancer Genomics, n.dcBioPortal for Cancer Genomics. Memorial Sloane Kettering Cancer Centre, www.cbioportal.org, (Accessed: 23/03/2016).Google Scholar) along with the recently published prostate TCGA dataset (Robinson et al., 2015Robinson D. Van Allen E.M. Wu Y.M. Schultz N. Lonigro R.J. Mosquera J.M. et al.Integrative clinical genomics of advanced prostate cancer.Cell. 2015; 161: 1215-1228Summary Full Text Full Text PDF PubMed Scopus (2011) Google Scholar). Here we introduce the camcAPP (http://bioinformatics.cruk.cam.ac.uk/apps/camcAPP/); a bespoke web interface to multiple prostate cancer genomics datasets. The interface was created with Shiny (https://www.rstudio.com/products/shiny/), and allows the non-specialist Bioinformatician to create publication-ready figures and tables through an intuitive interface to the underlying computer code. After selecting a dataset of interest, and uploading a list of genes, the following analyses can be performed:1)Boxplots and analysis of variance for expression of genes of interest grouped by clinical group, sample type, Gleason grade of copy-number cluster (N.B. not all covariates available for all data sets) (see Supplementary Fig. 2).2)A recursive partitioning-based survival analysis and Kaplan-Meier plots on a gene-by-gene basis (Supplementary Fig. 3).3)Pairwise-correlations of gene expression across whole studies and within clinical subgroups.4)Clustering and heatmaps of gene expression data, with options to interrogate associations between clinical covariates and newly-derived clusters (Supplementary Fig. 4).5)Tabulating the number of copy-number amplifications and deletions observed across whole studies or within a particular clinical covariate, and making a heatmap of copy-number calls (Supplementary Fig. 5) (Lalonde et al., 2014Lalonde E. Ishkanian A.S. Sykes J. Fraser M. Ross-Adams H. Erho N. et al.Tumour genomic and microenvironmental heterogeneity for integrated prediction of 5-year biochemical recurrence of prostate cancer: a retrospective cohort study.Lancet Oncol. 2014; 15: 1521-1532Summary Full Text Full Text PDF PubMed Scopus (236) Google Scholar). One of the challenges in constructing such a tool is delivering an output format that is readily transferable to slides for presentations or panels of a figure for publication. We recognise that this is, in part, a matter of axis typesetting and plot configuration but also of delivering an output file which permits further adjustment of the figure in, for example, Adobe Illustrator™. To this end, all plots can be exported as PDF or PNG files with configurable dimensions. Furthermore, for those that are well-versed in R, the code to produce a particular plot can be downloaded and modified as required. A further challenge that we seek to address with this interface is merging datasets for combined analysis. We hope to offer this option in due course, as we include further datasets that include samples analysed on compatible platforms. Strategies to address the Big Data problem have focussed on making the ever-increasing volume of genomic data accessible to scientists and on opening up the possibility of engaging non-specialists (Keener, 2015Keener Amanda B. The Scientist. July 8, 2015; (Accessed: 23/3/16)http://www.the-scientist.com/?articles.view/articleNo/43483/title/Big-Data-Problem/Google Scholar). This approach embodies a responsible attitude to science both in terms of patient input and financial resource and we believe that tools such as this are an important step to maximising the value of these landmark studies. We take pleasure in making this platform available to the prostate cancer community by means of this ‘In Focus’ article in EBioMedicine, a journal that we believe champions this responsible approach to genomic data both in cancer genomics (Kerns et al., 2016Kerns S.L. Dorling L. Fachal L. Bentzen S. Pharoah P.D. Barnes D.R. et al.Meta-analysis of genome wide association studies identifies genetic markers of late toxicity following radiotherapy for prostate cancer: radiogenomics consortium.EBioMedicine. 2016; 10: 150-163https://doi.org/10.1016/j.ebiom.2016.07.022Summary Full Text Full Text PDF PubMed Scopus (56) Google Scholar) and further afield (Taudien et al., 2016Taudien S. Lausser L. Giamarellos-Bourboulis E.J. Sponholz C. Schöneweck F. Felder M. et al.Genetic factors of the disease course after sepsis: rare deleterious variants are predictive.EBioMedicine. 2016; 12: 227-238https://doi.org/10.1016/j.ebiom.2016.08.037Summary Full Text Full Text PDF PubMed Scopus (27) Google Scholar). Core CRUK funding: MD, AGL, ADL.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,001 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».