Associations of breast cancer etiologic factors with stromal microenvironment of primary invasive breast cancers in the Ghana Breast Health Study
Bibliographic record
Abstract
Abstract Background: Emerging data suggest that beyond the neoplastic parenchyma, the stromal microenvironment (SME) impacts tumor biology, including aggressiveness, metastatic potential, and response to treatment. However, the epidemiological determinants of SME biology remain poorly understood, more so among women of African ancestry who are disproportionately affected by aggressive breast cancer phenotypes. Methods: Within the Ghana Breast Health Study, a population-based case-control study in Ghana, we applied high-accuracy machine-learning algorithms to characterize biologically-relevant SME phenotypes, including tumor-stroma ratio (TSR (%); a metric of connective tissue stroma to tumor ratio) and tumor-associated stromal cellular density (Ta-SCD (%); a tissue biomarker that is reminiscent of chronic inflammation and wound repair response in breast cancer), on digitized H&E-stained sections from 792 breast cancer patients aged 17–84 years. Kruskal-Wallis tests and multivariable linear regression models were used to test associations between established breast cancer risk factors, tumor characteristics, and SME phenotypes. Results: Decreasing TSR and increasing Ta-SCD were strongly associated with aggressive, mostly high grade tumors (p-value < 0.001). Several etiologic factors were associated with Ta-SCD, but not TSR. Compared with nulliparous women [mean (standard deviation) = 28.9% (7.1%)], parous women [mean (standard deviation) = 31.3% (7.6%)] had statistically significantly higher levels of Ta-SCD (p-value = 0.01). Similarly, women with a positive family history of breast cancer [FHBC; mean (standard deviation) = 33.0% (7.5%)] had higher levels of Ta-SCD than those with no FHBC [mean (standard deviation) = 30.9% (7.6%); p-value = 0.01]. Conversely, increasing body size was associated with decreasing Ta-SCD [mean (standard deviation) = 32.0% (7.4%), 31.3% (7.3%), and 29.0% (8.0%) for slight, moderate, and large body sizes, respectively, p-value = 0.005]. These associations persisted and remained statistically significantly associated with Ta-SCD in mutually-adjusted multivariable linear regression models (p-value < 0.05). With the exception of body size, which was differentially associated with Ta-SCD by grade levels (p-heterogeneity = 0.04), associations between risk factors and Ta-SCD were not modified by tumor characteristics. Conclusions: Our findings raise the possibility that epidemiological factors may act via the SME to impact both risk and biology of breast cancers in this population, underscoring the need for more population-based research into the role of SME in multi-state breast carcinogenesis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".