Abstract 2607: The cBioPortal for Cancer Genomics: an open source platform for accessing and interpreting complex cancer genomics data in the era of precision medicine
Bibliographic record
Abstract
Abstract The cBioPortal for Cancer Genomics is an open-access portal (http://cbioportal.org) that enables interactive, exploratory analysis of large-scale cancer genomics data. It integrates genomic and clinical data, and provides a suite of visualization and analysis options, including cohort and patient-level visualization, mutation visualization, survival analysis, enrichment analysis, and network analysis. The user interface is user-friendly, responsive, and makes genomic data easily accessible to translational scientists, biologists, and clinicians. The cBioPortal is a fully open source platform. All code is available on GitHub (https://github.com/cBioPortal/) under GNU Affero GPL license. The code base is maintained by multiple groups, including Memorial Sloan Kettering Cancer Center, Dana-Farber Cancer Institute, Children’s Hospital of Philadelphia, Princess Margaret Cancer Centre, and The Hyve, an open source bioinformatics company based in the Netherlands. More than 30 academic centers as well as multiple pharmaceutical and biotech companies maintain private instances of the cBioPortal. This includes the recently launched cBioPortal instance at the NCI Genomic Data Commons (https://cbioportal.gdc.nci.nih.gov/), and two large cBioPortal instances hosting genomic and clinical data at MSK and DFCI, supporting the MSK-IMPACT and DFCI Profile projects, two of the largest clinical sequencing efforts in the world. Our multi-institutional software team has accelerated the progress of evolving the core architectural technologies and developing new features to keep pace with the rapidly advancing fields of cancer genomics and precision cancer medicine. For example, we have integrated multi-platform genomics data with extensive clinical data including patient demographics, treatment history, and survival data. We have also developed a patient-centric view that visualizes both clinical and genomic data with annotation from OncoKB knowledge base. In the next few years, the development team will focus on the following areas: (1) Implementing major architectural changes to ensure future scalability and performance. (2) New features to support precision medicine, including (i) improved integration of knowledge base annotation, (ii) enhanced visualization of patient timeline, drug response, and tumor evolution, (iii) new patient similarity metrics, (iv) improved support for immunogenomics and immunotherapy, and (v) new visualization and analysis features for understanding response to therapy. (3) New analysis and target discovery features for large cohorts, including (i) supporting user-defined virtual cohort by selecting samples from multiple studies, and (ii) comparison of genomic or clinical characteristics of two or more selected cohorts. (4) Expanding community outreach, user support and training, and documentation. Citation Format: Jianjiong Gao, Ersin Ciftci, Pichai Raman, Pieter Lukasse, Istemi Bahceci, Adam Abeshouse, Hsiao-Wei Chen, Ino de Bruijn, Benjamin Gross, Zachary Heins, Ritika Kundra, Aaron Lisman, Angelica Ochoa, Robert Sheridan, Onur Sumer, Yichao Sun, Jiaojiao Wang, Manda Wilson, Hongxin Zhang, James Xu, Andy Dufilie, Priti Kumari, James Lindsay, Anthony Cros, Karthik Kalletla, Fedde Schaeffer, Sander Tan, Sjoerd van Hagen, Jorge Reis-Filho, Kees van Bochove, Ugur Dogrusoz, Trevor Pugh, Adam Resnick, Chris Sander, Ethan Cerami, Nikolaus Schultz. The cBioPortal for Cancer Genomics: an open source platform for accessing and interpreting complex cancer genomics data in the era of precision medicine [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 2607. doi:10.1158/1538-7445.AM2017-2607
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.015 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.005 | 0.004 |
| Open science | 0.006 | 0.010 |
| Research integrity | 0.003 | 0.005 |
| Insufficient payload (model declined to judge) | 0.064 | 0.078 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".