MétaCan
Menu
Back to cohort
Record W2740384655 · doi:10.1158/1538-7445.am2017-378

Abstract 378: The Cancer Genome Collaboratory

2017· article· en· W2740384655 on OpenAlexaff
Christina K. Yung, George L. Mihaiescu, Bob Tiernay, Junjun Zhang, Francois Gerthoffert, Andy Yang, Jared Baker, Guillaume Bourque, Paul C. Boutros, Bartha Maria Knoppers, B. F. Francis Ouellette, Cenk Sahinalp, Sohrab P. Shah, Vincent Ferretti, Lincoln Stein

Bibliographic record

VenueCancer Research · 2017
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsSimon Fraser UniversityConcordia UniversityMcGill UniversityBC Cancer AgencyOntario Institute for Cancer Research
Fundersnot available
KeywordsCollaboratoryCloud computingInteroperabilityComputer scienceGenomeMetadataSoftwareGenomicsComputational biologyData miningData scienceBiologyWorld Wide WebGeneticsGeneOperating system

Abstract

fetched live from OpenAlex

Abstract The Cancer Genome Collaboratory is an academic compute cloud designed to enable computational research on the world’s largest and most comprehensive cancer genome dataset, the International Cancer Genome Consortium (ICGC). The ICGC is on target to categorize the genomes of 25,000 tumors by 2018. A subproject of ICGC, the PanCancer Analysis of Whole Genomes (PCAWG) alone has generated over 800TB of harmonized sequence alignments, variants and interpreted data from over 2,800 cancer patients. A dataset of this size requires months to download and significant resources to store and process. By making the ICGC data available in cloud compute form in the Collaboratory, researchers can bring their analysis methods to the cloud, yielding benefits from the high availability, scalability and economy offered by cloud services, avoiding a large investment in static compute resources and essentially eliminating the time needed to download the data. To facilitate the computational analysis on the ICGC data, the Collaboratory has developed software solutions that are optimized for typical cancer genomics workloads, including well tested and accurate genome aligners and somatic variant calling pipelines. We have developed a simple to use, but fast and secure, data transfer tool that imports genomic data from cloud object storage into the user’s compute instances. Because a growing number of cancer datasets have restrictions on their storage locations, it is important to have software solutions that are interoperable across multiple cloud environments. We have successfully demonstrated interoperability across The Cancer Genome Atlas (TCGA) dataset hosted at University of Chicago’s Bionimbus Protected Data Cloud, the ICGC dataset hosted at the Collaboratory, and ICGC datasets stored in the Amazon Web Services (AWS) S3 storage. Lastly, we have developed a non-intrusive user authorization system that allows the Collaboratory to authenticate against the ICGC Data Access Compliance Office (DACO) when researchers require access to controlled tier data. We anticipate that our software solutions will be implemented on additional commercial and academic clouds. The Collaboratory is actively growing, with a target hardware infrastructure of over 3000 CPU cores and 15 petabytes of raw storage. As of November 2016, the Collaboratory holds information on 2,000 ICGC PCAWG donors (500TB total). We anticipate expanding the Collaboratory to host the entire ICGC dataset of 25,000 donors (approximately 5PB) and to extend its data management and analysis facilities across multiple clouds. During the current closed beta phase, the Collaboratory has been successfully utilized by multiple research groups, most notably PCAWG project researchers who analyzed thousands of genomes at scale over a few weeks’ time. The Collaboratory will open to the public during the second quarter of 2017. We invite cancer researchers to learn more about our cloud resources at cancercollaboratory.org, and apply for access to the Collaboratory. Citation Format: Christina K. Yung, George L. Mihaiescu, Bob Tiernay, Junjun Zhang, Francois Gerthoffert, Andy Yang, Jared Baker, Guillaume Bourque, Paul C. Boutros, Bartha M. Knoppers, BF Francis Ouellette, Cenk Sahinalp, Sohrab P. Shah, Vincent Ferretti, Lincoln D. Stein. The Cancer Genome Collaboratory [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 378. doi:10.1158/1538-7445.AM2017-378

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.015
metaresearch head score (Gemma)0.018
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Other · Consensus signal: none
Teacher disagreement score0.116
Threshold uncertainty score0.389

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0150.018
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.005
Science and technology studies0.0030.001
Scholarly communication0.0060.003
Open science0.0040.008
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.1160.058

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.067
GPT teacher head0.416
Teacher spread0.349 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreOther

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2017
Admission routes1
Has abstractyes

Explore more

Same venueCancer ResearchSame topicCancer Genomics and DiagnosticsFrench-language works237,207