MétaCan
Menu
Back to cohort

Abstract B033: Identifying triple-negative breast cancer patients at high risk of worse prognosis using molecular features derived from histology images

2025· article· en· W4412163871 on OpenAlexaboutno aff
Jung Hun Oh, Fresia Pareja, Rena Elkin, Larry Norton, Joseph O. Deasy

Bibliographic record

VenueClinical Cancer Research · 2025
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicMolecular Biology Techniques and Applications
Canadian institutionsnot available
Fundersnot available
KeywordsHistologyMedicineBreast cancerTriple-negative breast cancerCancerTriple negativePathologyOncologyInternal medicine

Abstract

fetched live from OpenAlex

Abstract Triple-negative breast cancer (TNBC) is the most aggressive subtype of invasive breast cancer characterized by the lack of estrogen receptor (ER), progesterone receptor (PR), and HER2 expression. TNBC is heterogeneous in terms of the biological and clinical perspective and a subset of these tumors exhibits markedly poor prognosis. Identification of this TNBC subset is an unmet clinical need. To identify patients with primary TNBC at high risk of worse prognosis, we developed a graph (network)-based analysis approach combined with an unbalanced optimal transport technique, utilizing molecular features derived from histology images. A total of 143 H&E-stained histology images from The Cancer Genome Atlas (TCGA) primary TNBC cases were analyzed. Tumor tissues were segmented on whole-slide images using a pre-trained ResNet18 model with a patch size of 512×512 and a processing resolution of 0.5 microns per pixel. A pre-trained ResNet34 model was then used to estimate four molecular features—microsatellite instability, hypermutation density, chromosomal instability, and TP53 mutation—on the segmented tumor tissues. In addition, the spatial fraction of tumor-infiltrating lymphocytes (TILs), derived from histology images (Saltz et al., Cell Reports, 2018), and four morphology features (epithelial area, tubule formation, nuclear pleomorphism, and mitosis; Thennavan et al., Cell Genomics, 2021) graded by the breast cancer pathology expert committee were analyzed. Following exclusion of cases with incomplete data, 113 cases were used for network analysis. A feature network was constructed using nine histology-derived features, based on Spearman’s correlation. K-means clustering, employing unbalanced optimal transport to calculate Wasserstein distance, was used to identify subgroups in the resulting feature network. The Wasserstein distance computed between samples and cluster centroids served as the cost function during the K-means clustering process. The two identified subgroups, categorized as a high-risk group (N=81) and a low-risk group (N=32) based on disease-specific survival (DSS) rates, showed a statistically significant difference in DSS (log-rank p=0.047). Estimated TILs were significantly different between the high and low-risk groups (p<0.0001). CIBERSORT scores that quantify 22 immune cell types were assessed. The low-risk group showed significantly higher CD8 T cells (p=0.030), regulatory Tregs T cells (p=0.029), and M1 macrophages (p=0.006), whereas the high-risk group showed significantly higher M0 macrophages (p=0.026) and M2 macrophages (p=0.006). Restricting the analysis to TNBC cases with tumor stage ≥2 revealed a greater DSS difference between the high (N=64) and low-risk (N=28) groups (p=0.026). CIBERSORT analysis revealed that the low-risk group had significantly higher levels of CD8 T cells (p=0.009) and activated CD4 memory T cells (p=0.027). Our analyses show that a cold immune milieu characterized by low TILs and a dominance of non-activated and anti-inflammatory macrophages are associated with poor prognosis in TNBC. Citation Format: Jung Hun Oh, Fresia Pareja, Rena Elkin, Larry Norton, Joseph Deasy. Identifying triple-negative breast cancer patients at high risk of worse prognosis using molecular features derived from histology images [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B033.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.001
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.009
Threshold uncertainty score0.019

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0000.001
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0020.001
Science and technology studies0.0000.000
Scholarly communication0.0010.000
Open science0.0000.000
Research integrity0.0010.000
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.063
GPT teacher head0.467
Teacher spread0.404 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueClinical Cancer ResearchSame topicMolecular Biology Techniques and ApplicationsFrench-language works237,207