Abstract B039: Detecting obesity-associated histopathology characteristics in breast cancer using an AI foundation model
Notice bibliographique
Résumé
Abstract Background: Obesity is a known risk factor for breast cancer, particularly in postmenopausal women. Obesity-induced changes, including inflamed adipocytes, may drive tumorigenesis, progression, and metastasis, leading to worse outcomes and reduced therapy response. We hypothesized that histologic patterns in tumor and stromal regions of obese breast cancer patients could serve as features for training AI models that classify BMI. We further hypothesized that these features are more pronounced in postmenopausal women with HR+/HER2- breast cancer. Methods: H&E-stained tissue samples and clinical data, including BMI, were obtained from the Murtha Cancer Center at Walter Reed National Military Medical Center. Samples were digitized with a VS200 Olympus scanner at 20X and split into 380×380-pixel tiles. AI-based filters removed non-tissue tiles, including those with adipose tissue. We used the CanvOI foundation model to generate slide-level embeddings and applied multiple ML classifiers to distinguish high and low BMI patients. Additionally, a CanvOI-based cancer detector isolated tumor and stroma regions, allowing predictions based on the entire tissue, tumor, or stroma regions. Patients with high BMI (≥30) were compared to those with BMI < 25. Patients were grouped into four cohorts: (1) all patients (n=195, 102 with high BMI), (2) ER+/HER2- patients (n=134, 64 high), (3) patients from cohort 1 aged over 50 for postmenopausal selection (n=144, 85 high), and (4) patients from cohort 2 aged over 50 (n=100, 56 high). Pre-menopausal cohorts were excluded due to insufficient cases. Each cohort was split into train and test sets (75/25) for model training and evaluation. Results: The models identified patients with a BMI ≥30, achieving AUC values of 0.71, 0.68, 0.63, and 0.68 across the four cohorts respectively when analyzing the entire tissue. When restricted to tumor regions, the AUC values shifted to 0.71, 0.61, 0.64, and 0.65, respectively, showing similar performance. Using only stromal regions resulted in AUC values of 0.78, 0.62, 0.57, and 0.75 showing a slight improvement for the all-patients group and the post-menopausal ER+/HER2 group. Conclusions: The results confirm the presence of obesity-related signatures in both tumor and stromal regions of H&E-stained tissue. The findings suggest that stromal regions exhibit stronger signatures, reflected by higher AUC values. Subcohort analysis found no clinical or demographic traits that would improve classification performance, suggesting that morphological features associated with high BMI are present in all groups. This study also highlights foundation models’ utility in extracting demographic variables from archival samples. DISCLAIMER: The contents of this publication are the sole responsibility of the author(s) and do not necessarily reflect the views, opinions or policies of USUHS, HJF, the DoD or the Departments of the Army, Navy or Air Force. Mention of trade names, commercial products, or organizations does not imply endorsement by the U.S. Government. Citation Format: Edwin A. Heredia, John Paine, Shiva Patre, Thomas Jonsson, Thomas Keller, Zihang Fang, Valerie Narumi, Irika Katiyar, Ian Lagerstrom, Jamie Lombardo, Jerry S. H. Lee, David B. Agus, Reva Basho, Naim Matasci. Detecting obesity-associated histopathology characteristics in breast cancer using an AI foundation model [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B039.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,002 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,001 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,001 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,002 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».