Accurate prediction of BRCA1 and BRCA2 heterozygous genotypes using expression profiling of lymphocytes after irradiation-induced DNA damage
Bibliographic record
Abstract
Germline mutations in BRCA1 and BRCA2 genes predispose women to an increased risk of breast/ovarian cancer. Both genes have important roles in DNA damage repair and are implicated in gene expression regulation. We have previously shown that normal fibroblasts from mutation carriers can be distinguished from noncarriers following radiation-induced DNA damage. In this new study we used lymphocytes to determine whether these also show differential response to induced DNA damage and whether expression profiling using microarray technology could be used to accurately predict the BRCA genotype. Short-term lymphocyte cultures were established from fresh blood samples from 20 BRCA1 and 20 BRCA2 mutation carriers and from 10 negative controls (individuals tested negative for the mutation present in the family). Lymphocytes were subjected to 8 Gy ionizing irradiation to induce DNA damage and RNA was extracted 1 hour post γ-irradiation. For expression profiling, genome-wide (30 K) spotted cDNA microarrays manufactured by the Cancer Research UK Microarray Facility were used. We then applied the support vector machine (SVM) classifier with statistical feature selection to determine the best feature set for predicting BRCA1 and BRCA2 heterozygous genotypes. We also investigated the prediction accuracy using a nonprobabilistic classifier (SVM) and a probabilistic classifier (Gaussian process classifier).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".