Abstract 4743: A population-based approach to address clinical cancer care: The national genomics platform
Bibliographic record
Abstract
Abstract Our understanding of the genetic basis for cancer is advancing at a rapid rate due to the application of next generation sequencing (NGS) technologies. It is becoming common to sequence tumors and patients in an attempt to find actionable mutations that could offer the patient a more targeted and effective treatment. However, the genomics of cancer is very complex and only a handful of actionable mutations have been characterized. Larger-scale studies are needed to understand the pathogenic spectrum of cancer variants and deliver reliable clinical decision-making support to providers. Several large-scale sequencing projects are underway to collect NGS data on tumors (e.g. TCGA, ICGC, and TARGET) but a national platform for analysis, interpretation, and reporting does not exist. We propose that a consolidated informatics platform for the collection of outcomes data with genomics and clinical data would accelerate research and provide patients with the opportunity for personalized cancer treatment. In addition, with 1,665,540 new cancer cases predicted in the US for 2014, a national-scale genomics platform is needed, capable of sequencing 3 million genomes per year, storing exabytes of data, while supporting over 20,000 oncologists, researchers, and analysts. At Lockheed Martin, we deliver highly scalable and reliable information systems for a variety of missions and citizen services. Here, we will present our vision for a national cancer genomics platform to include NGS data collection, cost-efficient storage, scalable and modular processing pipelines, and collaborative analytics and data sharing capabilities, within a compliant privacy and security framework. In conclusion, by leveraging the scale of clinical cancer sequencing and capturing these data into a case management system for translational research, this platform provides data at-scale needed for finding actionable mutations, designing effective treatments and implementing prevention strategies, faster. Citation Format: Ogan Abaan, Amrita Basu, Noah Brown, Bret Light, David Deal, Michael Hultner. A population-based approach to address clinical cancer care: The national genomics platform. [abstract]. In: Proceedings of the 106th Annual Meeting of the American Association for Cancer Research; 2015 Apr 18-22; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2015;75(15 Suppl):Abstract nr 4743. doi:10.1158/1538-7445.AM2015-4743
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.018 | 0.021 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.006 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.026 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".