Multi-tissue transcriptome-wide association studies identified 235 genes for intrinsic subtypes of breast cancer
Bibliographic record
Abstract
BACKGROUND: Although genome-wide association studies (GWAS) of breast cancer (BC) identified common variants which differ between intrinsic subtypes, genes through which these variants act to impact BC risk have not been fully established. Transcriptome-wide association studies (TWAS) have identified genes associated with overall BC risk, but subtype-specific differences are largely unknown. METHODS: We performed two multi-tissue TWAS for each BC intrinsic subtype, including an expression-based approach that collated TWAS signals from expression quantitative trait loci (eQTLs) across multiple tissues and a novel splicing-based approach that collated signals from splicing QTLs (sQTLs) across intron clusters and subsequently across tissues. We used summary statistics for five intrinsic subtypes including Luminal A-like, Luminal B-like, Luminal B/HER2-negative-like, HER2-enriched-like, and triple-negative BC, generated from 106 278 BC cases and 91 477 controls in the Breast Cancer Association Consortium. RESULTS: Overall, we identified 235 genes in 88 loci that were associated with at least one of the five intrinsic subtypes. Most genes were subtype-specific, and many have not been reported in previous TWAS. We discovered common variants that modulate expression of CHEK2 confer increased risk to Luminal A-like BC, in contrast to the viewpoint that CHEK2 primarily harbors rare, penetrant mutations. Additionally, our splicing-based TWAS provided population-level support for MDM4 splice variants that increased the risk of triple-negative BC. CONCLUSION: Our comprehensive, multi-tissue TWAS corroborated previous GWAS loci for overall BC risk and intrinsic subtypes, while underscoring how common variation that impacts expression and splicing of genes in multiple tissue types can be used to further elucidate the etiology of BC.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".