Metadata record for the article: Current gene panels account for nearly all homologous recombination repair-associated multiple-case breast cancer families
Bibliographic record
Abstract
<b>Summary</b><br> This metadata record provides details of the data supporting the claims of the related article: “Current gene panels account for nearly all homologous recombination repair-associated multiple-case breast cancer families”. The related study hypothesised that variants in underexplored homologous recombination repair (HR) genes could explain unsolved multiple-case breast cancer (BC) families, and investigated HR deficiency (HRD)-associated mutational signatures and second hits in tumour DNA from familial BC cases. Type of data: tumour and germline data; whole exome sequencing Subject of data: <i>Homo sapiens</i> Sample size: 38 Population characteristics: Candidate families with early-onset BC or a strong family history of breast cancer. Blood samples were obtained directly from the participants at the time they were consented into the biobank or study. Tumour blocks were requested from pathology archives at various hospitals served by the CHUM and McGill genetic services in the greater western Quebec region. <b>Data access</b> Tumour and germline datasets generated during the current study are not publicly available as they are old samples that precede the reporting standards requiring deposition in public repositories, and, therefore, the consent provided by study participants did not include a provision for widespread disclosure. However, direct requests can be made to the corresponding authors for access. The following files are openly available as part of this <i>figshare </i>data record: ‘Master_list_variant.xlsx’ (underlying Table 1), ‘Genes assessed initially.xlsx’ (underlying Supplementary Table 1), ‘Candidate Genes list.xlsx’ (underlying Supplementary Table 2). All other data are housed on institutional storage and are not openly available in order to protect patient privacy as informed consent to share participant-level data was not obtained prior to or during data collection. However, these data can be requested from Dr Foulkes. The files are: ‘all germline_sample.vcf’, ‘data.tumor.maf’, ‘all tumor_sample.vcf’, ‘all sample.out.gzsmall.seqz.gz’. <b>Corresponding author(s) for this study</b> William D Foulkes, 1Department of Human Genetics, McGill University, 3640 Rue University, Room W-315D, Montreal, QC, H3A 0C7, Canada. william.foulkes@mcgill.ca Paz Polak. 13Department of Oncological Sciences, Icahn School of Medicine at Mount Sinai Hospital, Gustave L. Levy Place, NY 10029-5674 New York, USA. paz.polak@mssm.edut<br> <br> <b></b> <b>Study approval </b> Participants to this study were consented at two Montreal sites. The majority of participants were consented into a biobank entitled “Banque d'échantillons biologiques et de données (cliniques et biologiques) associées à des fins de recherche sur les cancers métastatiques du sein et de l'ovaire” created in 2000 and approved by the CHUM Institutional Review Board, approval number BD 04.002. A few additional participants were consented directly into the linked scientific study created and approved by the McGill University Institution Review Board in 2011 called “Genome-wide approaches in hereditary cancer families”, study number A08-M61-09B.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".