Abstract 1117: cBioPortal for cancer genomics
Bibliographic record
Abstract
Abstract cBioPortal for Cancer Genomics is a widely used platform for exploratory, interactive visualization and analysis of large-scale clinico-genomic datasets. cBioPortal provides a range of visualizations and analyses including interactive cohort exploration, OncoPrints, mutation “lollipop” plots, survival analysis, alteration enrichment analysis, and detailed patient-level visualizations. cBioPortal also integrates variant annotations from a variety of sources to facilitate interpretation. The public cBioPortal (https://www.cbioportal.org) is accessed by >40,000 unique visitors each month and hosts data from >460 studies. All data is also available in the cBioPortal Datahub: https://github.com/cBioPortal/datahub. In 2024 we added 76 new studies (∼30,000 samples), including data from the NCI Genomic Data Commons. In addition, >94 instances of cBioPortal are installed at academic institutions and companies worldwide. cBioPortal partners with AACR Project GENIE to provide access to the GENIE cohort in a dedicated instance (https://genie.cbioportal.org). Users can explore the full GENIE cohort of >229,000 clinically sequenced samples from 19 institutions, as well as cohorts with comprehensive clinical annotations including response, outcome, and treatment history, from the GENIE Biopharma Collaborative (BPC). BPC cohorts for NSCLC (∼2,000 samples) and colorectal cancer (∼1,500 samples) are available, with more to come. The past year has brought a variety of enhancements to cBioPortal. A new data type selector on the home page enables users to find studies with specific types of data. The interactive cohort exploration has new ways to explore data with the addition of gene-specific charts to summarize the types of mutations in a gene and the integration of the Plots tab for customizable graphs of any two data attributes. The OncoPrint can now display per group alteration frequency based on any categorical attribute. Variant interpretation is enhanced with the integration of AlphaMissense as a novel annotation source and an update to the latest MutationAssessor data. The patient page also has new visualizations, including mutational signatures and the integration of Chromoscope to visualize structural variations. We also made significant changes to the backend code to improve both the developer and user experience. The backend code was repackaged and upgraded to simplify and improve the development process. In addition, we are working on switching to an Online Analytical Processing (OLAP) database which will bring significant performance improvements. cBioPortal is open source: https://github.com/cBioPortal. Development is a collaborative effort among groups at Memorial Sloan Kettering Cancer Center, Dana-Farber Cancer Institute, Children’s Hospital of Philadelphia, Princess Margaret Cancer Centre, Caris Life Sciences, Bilkent University, SE4BIO and The Hyve. We welcome open source contributions from others in the cancer research community. Citation Format: Ino de Bruijn,Tali Mazor,Rima AlHamad,Calla Chennault,Corey Dubin,Jeremy Easton-Marks,Zhaoyuan Fu,Benjamin Gross,Charles Haynes,David M. Higgins,Jason Hwee,Prasanna K. Jagannathan,Mirella Kalafati,Karthik Kalletla,Zeynep Karagöz,James Ko,Tim Kuijpers,Sowmiyaa Kumar,Priti Kumari,Ritika Kundra,Bryan Lai,Xiang Li,James Lindsay,Aaron Lisman,Qi-Xuan Lu,Ramyasree Madupuri,Zain-ul-Abideen Nasir,Angelica Ochoa,Yusuf Ziya Özgül,Oleguer Plantalech,Matthijs N. Pon,Baby A. Satravada,Jessica Singh,Selcuk Onur Sumer,Pim van Nierop,Floris Vleugels,Avery Wang,Manda Wilson,Hongxin Zhang,Gaofei Zhao,Ugur Dogrusoz,Allison Heath,Adam Resnick,Trevor J. Pugh,Chris Sander,Ethan Cerami,JianJiong Gao,Nikolaus Schultz. cBioPortal for cancer genomics [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1117.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".