MétaCan
Menu
Back to cohort
Record W4402287509 · doi:10.1101/2024.09.04.24313051

Analysis of more than 400,000 women provides case-control evidence for BRCA1 and BRCA2 variant classification

2024· preprint· en· W4402287509 on OpenAlexaff
Maria Zanti, Denise G O'Mahony, Michael T. Parsons, Leila Dorling, Joe Dennis, Nicholas Boddicker, Wenan Chen, Chunling Hu, Marc Naven, Kristia Yiangou, Thomas U. Ahearn, Christine B. Ambrosone, Irene L. Andrulis, Antonis C. Antoniou, Paul L. Auer, Caroline Baynes, Clara Bodelón, Natalia Bogdanova, Stig E. Bojesen, Manjeet K. Bolla, Kristen D. Brantley, Nicola J. Camp, Archie Campbell, Jose E. Castelao, Melissa H. Cessna, Jenny Chang‐Claude, Fei Chen, Georgia Chenevix‐Trench, Don Conroy, Kamila Czene, Arcangela De Nicolo, Susan M. Domchek, Thilo Dörk, Alison M. Dunning, A. Heather Eliassen, D Gareth Evans, Peter A. Fasching, Jonine D. Figueroa, Henrik Flyger, Manuela Gago-Domínguez, Montserrat García‐Closas, Gord Glendon, Anna González‐Neira, Felix Graßmann, Andreas Hadjisavvas, Christopher A. Haiman, U. Hamann, Steven N. Hart, Mikael Hartman, Weang-Kee Ho, James M. Hodge, Reiner Hoppe, Sacha J. Howell, Anna Jakubowska, Elza K. Khusnutdinova, Yon-Dschun Ko, Peter Kraft, Vessela N. Kristensen, James V. Lacey, Jingmei Li, Geok Hoon Lim, Sara Lindström, Artitaya Lophatananon, Craig Luccarini, Arto Mannermaa, Maria Elena Martinez, Dimitrios Mavroudis, Roger L. Milne, Kenneth Muir, Katherine L. Nathanson, Rocío Núñez‐Torres, Nadia Obi, Janet E. Olson, Julie R. Palmer, Mihalis I. Panayiotidis, Alpa V. Patel, Paul D.P. Pharoah, Eric C. Polley, Muhammad Usman Rashid, Kathryn J. Ruddy, Emmanouil Saloustros, Elinor J. Sawyer, Marjanka K. Schmidt, Melissa C. Southey, Veronique Kiak‐Mien Tan, Soo‐Hwang Teo, Lauren R. Teras, Diana Torres, Amy Trentham‐Dietz, Thérèse Truong, Celine M. Vachon, Qin Wang, Jeffrey N. Weitzel, Siddhartha Yadav, Song Yao, Gary Zirpoli, Melissa Cline, Peter Devilee, Sean V. Tavtigian, David E Goldgar, Fergus J. Couch, Douglas F. Easton, Amanda B. Spurdle, Kyriaki Michailidou

Bibliographic record

VenuemedRxiv · 2024
Typepreprint
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicGenomics and Rare Diseases
Canadian institutionsLunenfeld-Tanenbaum Research InstituteUniversity of TorontoMount Sinai Hospital
FundersNational Institutes of HealthEuropean CommissionWellcome TrustBreast Cancer Research Foundation
KeywordsGermlineOdds ratioAlleleGenetic testingGermline mutationGeneticsComputational biologyMedicineBiologyMutationInternal medicineGene

Abstract

fetched live from OpenAlex

Abstract Clinical genetic testing identifies variants causal for hereditary cancer, information that is used for risk assessment and clinical management. Unfortunately, some variants identified are of uncertain clinical significance (VUS), complicating patient management. Case-control data is one evidence type used to classify VUS, and previous findings indicate that case-control likelihood ratios (LRs) outperform odds ratios for variant classification. As an initiative of the Evidence-based Network for the Interpretation of Germline Mutant Alleles (ENIGMA) Analytical Working Group we analyzed germline sequencing data of BRCA1 and BRCA2 from 96,691 female breast cancer cases and 303,925 unaffected controls from three studies: the BRIDGES study of the Breast Cancer Association Consortium, the Cancer Risk Estimates Related to Susceptibility consortium, and the UK Biobank. We observed 11,227 BRCA1 and BRCA2 variants, with 6,921 being coding, covering 23.4% of BRCA1 and BRCA2 VUS in ClinVar and 19.2% of ClinVar curated (likely) benign or pathogenic variants. Case-control LR evidence was highly consistent with ClinVar assertions for (likely) benign or pathogenic variants; exhibiting 99.1% sensitivity and 95.4% specificity for BRCA1 and 92.2% sensitivity and 86.6% specificity for BRCA2 . This approach provides case-control evidence for 785 unclassified variants, that can serve as a valuable element for clinical classification.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.011
metaresearch head score (Gemma)0.020
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.011
Threshold uncertainty score0.058

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0110.020
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.002
Science and technology studies0.0010.001
Scholarly communication0.0010.000
Open science0.0010.001
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0040.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.029
GPT teacher head0.299
Teacher spread0.270 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2024
Admission routes1
Has abstractyes

Explore more

Same venuemedRxivSame topicGenomics and Rare DiseasesFrench-language works237,207