Identification of cell proliferation, immune response and cell migration as critical pathways in a prognostic signature for HER2+:ERα- breast cancer
Bibliographic record
Abstract
BACKGROUND: Multi-gene prognostic signatures derived from primary tumor biopsies can guide clinicians in designing an appropriate course of treatment. Identifying genes and pathways most essential to a signature performance may facilitate clinical application, provide insights into cancer progression, and uncover potentially new therapeutic targets. We previously developed a 17-gene prognostic signature (HTICS) for HER2+:ERα- breast cancer patients, using genes that are differentially expressed in tumor initiating cells (TICs) versus non-TICs from MMTV-Her2/neu mammary tumors. Here we probed the pathways and genes that underlie the prognostic power of HTICS. METHODS: We used Leave-One Out, Data Combination Test, Gene Set Enrichment Analysis (GSEA), Correlation and Substitution analyses together with Receiver Operating Characteristic (ROC) and Kaplan-Meier survival analysis to identify critical biological pathways within HTICS. Publically available cohorts with gene expression and clinical outcome were used to assess prognosis. NanoString technology was used to detect gene expression in formalin-fixed paraffin embedded (FFPE) tissues. RESULTS: We show that three major biological pathways: cell proliferation, immune response, and cell migration, drive the prognostic power of HTICS, which is further tuned by Homeostatic and Glycan metabolic signalling. A 6-gene minimal Core that retained a significant prognostic power, albeit less than HTICS, also comprised the proliferation/immune/migration pathways. Finally, we developed NanoString probes that could detect expression of HTICS genes and their substitutions in FFPE samples. CONCLUSION: Our results demonstrate that the prognostic power of a signature is driven by the biological processes it monitors, identify cell proliferation, immune response and cell migration as critical pathways for HER2+:ERα- cancer progression, and defines substitutes and Core genes that should facilitate clinical application of HTICS.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".