Integrated molecular profiles of invasive breast tumors and ductal carcinoma in situ (DCIS) reveal differential vascular and interleukin signaling
Bibliographic record
Abstract
We use an integrated approach to understand breast cancer heterogeneity by modeling mRNA, copy number alterations, microRNAs, and methylation in a pathway context utilizing the pathway recognition algorithm using data integration on genomic models (PARADIGM). We demonstrate that combining mRNA expression and DNA copy number classified the patients in groups that provide the best predictive value with respect to prognosis and identified key molecular and stromal signatures. A chronic inflammatory signature, which promotes the development and/or progression of various epithelial tumors, is uniformly present in all breast cancers. We further demonstrate that within the adaptive immune lineage, the strongest predictor of good outcome is the acquisition of a gene signature that favors a high T-helper 1 (Th1)/cytotoxic T-lymphocyte response at the expense of Th2-driven humoral immunity. Patients who have breast cancer with a basal HER2-negative molecular profile (PDGM2) are characterized by high expression of protumorigenic Th2/humoral-related genes (24-38%) and a low Th1/Th2 ratio. The luminal molecular subtypes are again differentiated by low or high FOXM1 and ERBB4 signaling. We show that the interleukin signaling profiles observed in invasive cancers are absent or weakly expressed in healthy tissue but already prominent in ductal carcinoma in situ, together with ECM and cell-cell adhesion regulating pathways. The most prominent difference between low and high mammographic density in healthy breast tissue by PARADIGM was that of STAT4 signaling. In conclusion, by means of a pathway-based modeling methodology (PARADIGM) integrating different layers of molecular data from whole-tumor samples, we demonstrate that we can stratify immune signatures that predict patient survival.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".