MétaCan
Menu
← Back to cohort
Record W2496068731 · doi:10.1158/1538-7445.am2016-850

Abstract 850: Comprehensive genome and transcriptome structural analysis of a breast cancer cell line using single molecule sequencing

2016· article· en· W2496068731 on OpenAlexaff
Maria Nattestad, Karen Ng, Sara Goodwin, Timour Baslan, Fritz J. Sedlazeck, James Gurtowski, Elizabeth R. Hutton, Yogi Sundaravadanam, Tyler H. Garvin, Marley C. Alford, Elizabeth Tseng, Philipp Rescheneder, Jason Chin, Timothy A. Beck, Melissa Kramer, John D. McPherson, James Hicks, Michael C. Schatz, W. Richard McCombie

Bibliographic record

VenueCancer Research · 2016
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsOntario Institute for Cancer Research
Fundersnot available
KeywordsChromothripsisStructural variationCopy number analysisBreakpointCopy-number variationGenomeContext (archaeology)Computational biologyBiologyGenome instabilityGeneticsCancerGenomicsBreast cancerCancer genome sequencingGeneChromosomeDNA

Abstract

fetched live from OpenAlex

Abstract Genomic instability is one of the hallmarks of cancer, leading to widespread copy number variations, chromosomal fusions, and other structural variations in many cancers. The breast cancer cell line SK-BR-3 is an important model for HER2+ breast cancers, which are among the most aggressive forms of the disease and affect one in five cases. Through short read sequencing, copy number arrays, and other technologies, the genome of SK-BR-3 is known to be highly rearranged with many copy number variations, including an approximately twenty-fold amplification of the HER2 oncogene, along with numerous other amplifications and deletions. However, these technologies cannot precisely characterize the nature and context of the identified genomic events and other important mutations may be missed altogether because of repeats, multi-mapping reads, and the failure to reliably anchor alignments to both sides of a variation. To address these challenges, we have sequenced SK-BR-3 using PacBio long read technology. Using the new P6-C4 chemistry, we generated more than 70X coverage of the genome with average read lengths of 9-13kb (max: 71kb). Using Lumpy for split-read alignment analysis, as well as our novel assembly-based algorithms for finding complex variants, we have developed a detailed map of structural variations in this cell line. Taking advantage of the newly identified breakpoints and combining these with copy number assignments, we have developed an algorithm to reconstruct the mutational history of this cancer genome. From this we have characterized the amplifications of the HER2 region, discovering a complex series of nested duplications and translocations between chr17 and chr8, two of the most frequent translocation partners in primary breast cancers. We have also carried out full-length transcriptome sequencing using PacBio's Iso-Seq technology, which has revealed a number of previously unrecognized gene fusions and isoforms. Combining long-read genome and transcriptome sequencing technologies enables an in-depth analysis of how changes in the genome affect the transcriptome, including how gene fusions are created across multiple chromosomes. This analysis has established the most complete cancer reference genome available to date, and is already opening the door to applying long-read sequencing to patient samples with complex genome structures. Citation Format: Maria Nattestad, Karen Ng, Sara Goodwin, Timour Baslan, Fritz Sedlazeck, James Gurtowski, Elizabeth Hutton, Yogi Sundaravadanam, Tyler Garvin, Marley Alford, Elizabeth Tseng, Philipp Rescheneder, Jason Chin, Timothy Beck, Melissa Kramer, John McPherson, James Hicks, Michael C. Schatz, William R. McCombie. Comprehensive genome and transcriptome structural analysis of a breast cancer cell line using single molecule sequencing. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 850.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.004
Threshold uncertainty score0.007

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.066
GPT teacher head0.359
Teacher spread0.293 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2016
Admission routes1
Has abstractyes

Explore more

Same venueCancer Research→Same topicCancer Genomics and Diagnostics→French-language works237,207→