Phenotypic characterization of mouse mammary epithelial stem and progenitor cells
Bibliographic record
Abstract
Background Epidemiological studies have shown that only about 20% of the familial clustering of breast cancer is explained by the known highly penetrant mutations in BRCA1 and BCRA2.We have set out to search for the genes for the remaining 80%.Twin studies indicate a predominant role of shared genes rather than a shared environment; the patterns of occurrence of breast cancer in families are consistent with a major polygenic component.Methods We have assembled a population based set of 5,000 breast cancer cases and 5,000 controls from the East Anglian population.We have simple clinical and epidemiological information, including family history, and samples of blood and paraffin embedded tumour.We have used association studies based on single nucleotide polymorphisms, first with candidate genes and then in a genome-wide scan of 266,000 single nucleotide polymorphisms, to search for the putative predisposing genes.We have as yet searched only for common variants (frequency >5%). ResultsWe have modelled the effects of polygenic predisposition in the East Anglian population, and have shown that the model predicts a wide distribution of individual risk in the population, such that half of all breast cancers may occur in the 12% of women at greatest risk.Both the candidate gene-based and genome-wide scans have provided provisional identification of a number of novel susceptibility genes, and these are currently being confirmed by a world-wide consortium of independent laboratories totalling 20,000-plus cases and controls.No single gene so far identified contributes more than 2% of the total inherited component, consistent with a model in which susceptibility is the result of a large number of individually small genetic effects. S2 Translating breast cancer research into clinical practice -new approaches and better outcomes
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".