Development of a Machine Learning-Based Prognostic Model for Hormone Receptor-Positive Breast Cancer Using Nine-Gene Expression Signature
Bibliographic record
Abstract
Background: Determining the prognosis of hormone receptor positive (HR + ) breast cancer (BC), which accounts for 80% of all BCs, is critical in improving survival outcomes. Stratifying individuals at high risk of BC-related mortality and improving prognosis has been the focus of research for over a decade. However, these tools are not universal as they are limited to clinical factors. We hypothesized that a new framework for predicting prognosis in HR + BC patients can develop using artificial intelligence. Methods: A total of 2,338 HR + human epidermal growth factor receptor 2 negative (HER2 - ) BC cases were analyzed from Molecular Taxonomy of Breast Cancer International Consortium (METABRIC), The Cancer Genome Atlas (TCGA), and Gene Expression Omnibus (GEO) cohorts. Groups were then divided into high- and low-risk categories utilizing a recurrence prediction model (RPM). An RPM was created by extracting nine prognosis-related genes from over 18,000 genes using a logistic progression model. Results: Risk classification by RPM was significantly stratified in both the discovery cohort and validation cohort. In the time-dependent area under the curve analysis, there was some variation depending on the cohort, but accuracy was found to decline significantly after about 10 years. Cell cycle related gene sets, MYC, and PI3K-AKT-mTOR signaling were enriched in high-risk tumors by the Gene Set Enrichment Analysis. High-risk tumors were associated with high levels of immune cells from the lymphoid and myeloid lineage and immune cytolytic activity, as well as low levels of stem cells and stromal cells. High-risk tumors were also associated with poor therapeutic effects of chemotherapy and endocrine therapy. Conclusions: This model was able to stratify prognosis in multiple cohorts. This is because the model reflects major BC therapeutic target pathways and tumor immune microenvironment and, further is supported by the therapeutic effect of chemotherapy and endocrine therapy. World J Oncol. 2023;14(5):406-422 doi: https://doi.org/10.14740/wjon1700
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".