MétaCan
Menu
← Back to cohort

Abstract PD8-01: The metastatic breast cancer project: Generating the clinical and genomic landscape of metastatic breast cancer through patient-partnered research

2020· article· en· W3013282680 on OpenAlexaboutno aff
Nikhil Wagle, Corrie Painter, Elana Anastasio, Michael Dunphy, Mary McGillicuddy, Esha Jain, Tania G. Hernandez, Sara Balch, Beena Thomas, Dewey Kim, Alyssa L. Damon, Shahrayz Shah, Brett N. Tomson, Rachel Stoddard, Colleen Nguyen, Jorge E. Buendia-Buendia, Ofir Cohen, Jorge Gómez Tejeda Zañudo, Netsanet Tsegai, Lauren Sterlin, Ulcha F. Ulysse, Kathryn Sine, Oyin Alao, Jacqueline Lucia, Eric S. Lander, Todd R. Golub

Bibliographic record

VenueCancer Research · 2020
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicCancer Genomics and Diagnostics
Canadian institutionsnot available
Fundersnot available
KeywordsBreast cancerMedicineCancerMetastatic breast cancerBiobankInternal medicineOncologyFamily medicineBioinformaticsBiology

Abstract

fetched live from OpenAlex

Abstract Background: The Metastatic Breast Cancer Project (MBCproject) is a research study that directly engages patients (pts) through social media and advocacy groups, and empowers them to share their samples, clinical information, and experiences. The goal is to create a publicly available dataset of linked genomic, clinical, and pt-reported data to enable research. Methods: In collaboration with pts, advocates, and advocacy groups, a website (MBCproject.org) was developed that allows pts with metastatic breast cancer (MBC) anywhere in the U.S. or Canada to register. Registered pts are sent an online consent form that asks for permission to obtain and analyze their medical records and samples. Once enrolled, pts are sent a saliva kit and a blood kit and asked to mail back a saliva sample, which is used to extract germline DNA, and/or a blood sample, which is used to extract germline DNA and cell free DNA (cfDNA). We contact participants’ medical providers and obtain medical records and a portion of their stored tumor biopsies. Whole exome sequencing (WES) is performed on tumor DNA, germline DNA, and cfDNA; transcriptome sequencing (RNA-seq) is performed on tumor RNA. Medical records and pt-reported data are abstracted to create a detailed clinical record for each pt. All de-identified data are shared regularly via public databases (cbioportal.org, mbcproject.org, dbGaP, NCI Genomic Data Commons) without restrictions. Study updates are shared with participants regularly. Results: From 10/20/15-7/8/19, 5357 women and men with MBC registered. 3290 pts receiving care at over 1700 different institutions consented to share medical records and tumor/saliva/blood samples, and to have genomic analysis performed. Details of clinical data collection, biospecimen acquisition, and genomic data generation to date are outlined in the Table. WES from 463 tumors obtained from 326 pts have been generated (with matched germline WES), including 61 pts with 2 timepoints, 19 pts w 3 timepoints, and 11 pts w 4+ timepoints. 278 tumor exomes were from the breast/regional lymph nodes, 63 from distant metastatic sites and 122 from cfDNA. 110 tumor exomes were from samples obtained before the diagnosis of MBC, 258 from after the diagnosis of MBC, and 95 to be determined (TBD). 161 tumor exomes were obtained prior to any therapy, 204 following some therapy, and 98 TBD. Clinically annotated genomic data are used to study specific pt cohorts (including rare subsets and outliers) and to identify mechanisms of response and resistance to therapies. Examples of the clinical and genomic analyses that will be presented include: - Pts diagnosed <40 yrs of age (1108 pts enrolled; 120 with tumor WES) - de novo MBC (1122 pts enrolled; 121 with tumor WES) - Late recurrence, >5 yrs after diagnosis (830 pts enrolled; 77 with tumor WES) - Long-term survivors, >10 yrs with MBC (159 pts enrolled; 11 with tumor WES) - Resistance to CDK4/6 inhibitors (709 pts enrolled; 148 with tumor WES) Conclusions: Partnering directly with pts enables rapid identification of thousands of pts willing to share tumors, blood, saliva, and medical records to accelerate research. This approach allows for identification of patients with specific phenotypes, who have been challenging to identify with traditional approaches. Remote acquisition of medical records and saliva/blood/tumor samples is feasible. This clinically annotated dataset is a shared resource for the research community. Table 1Clinical data collection, biospecimen acquisition, and genomic data generation:NumberConsent signed3290 ptsSurvey #1 submitted3290 pts(demographics, diagnosis details, receptor status, clinical experiences)Survey #2 submitted1435 pts(pathology details, sites of metastasis, treatments with start and stop dates)Medical record received1307 ptsSaliva sample received1976 ptsBlood sample received1121 ptsTumor samples received482 tumor samples from 346 ptsDigital image of tumor slide H&E generated482 tumor samplesWES from germline complete310 germline samplesWES from tumor sample complete341 tumor samplesRNA-seq from tumor sample complete229 tumor samplesULP-WGS from cfDNA complete947 blood samplesWES from circulating tumor DNA complete122 blood samples Citation Format: Nikhil Wagle, Corrie Painter, Elana Anastasio, Michael Dunphy, Mary McGillicuddy, Esha Jain, Tania G Hernandez, Sara Balch, Beena Thomas, Dewey Kim, Alyssa L. Damon, Shahrayz Shah, Brett N. Tomson, Rachel Stoddard, Colleen Nguyen, Jorge Buendia-Buendia, Ofir Cohen, Jorge Gomez Tejeda Zanudo, Netsanet Tsegai, Lauren Sterlin, Ulcha Fergie Ulysse, Kathryn Sine, Oyin Alao, Jacqueline Lucia, Eric S. Lander, Todd R. Golub. The metastatic breast cancer project: Generating the clinical and genomic landscape of metastatic breast cancer through patient-partnered research [abstract]. In: Proceedings of the 2019 San Antonio Breast Cancer Symposium; 2019 Dec 10-14; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2020;80(4 Suppl):Abstract nr PD8-01.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.017
metaresearch head score (Gemma)0.038
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.025
Threshold uncertainty score0.090

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0170.038
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.003
Science and technology studies0.0020.001
Scholarly communication0.0040.002
Open science0.0030.011
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0250.008

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.174
GPT teacher head0.456
Teacher spread0.282 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueCancer Research→Same topicCancer Genomics and Diagnostics→French-language works237,207→