Abstract B038: Machine learning used to validate neutrophil classification in Triple Negative Breast Cancer patients
Bibliographic record
Abstract
Abstract Introduction: Neutrophils drive triple-negative breast cancer (TNBC) progression through angiogenesis, immune suppression, and metastasis via neutrophil extracellular traps (NETs). Their prominence in TNBC’s tumor microenvironment and circulation positions them as potential biomarkers for tracking recurrence and immunotherapy response. Yet, their low RNA content and fragility complicate single-cell RNA-sequencing (scRNA-seq) within peripheral blood mononuclear cell (PBMC)-focused workflows, which inadvertently capture these non-PBMC cells. Materials and methods: We developed a bioinformatic tool centered on marker-based annotation to identify neutrophils from scRNA-seq data of peripheral blood from TNBC patients (N=7, 44 485 cells, yielding 45 neutrophils) and healthy controls (N=4, 35 391 cells, yielding 40 neutrophils). A machine learning framework based on XGBoost, trained on public datasets, was integrated to enhance annotation reliability, addressing marker overlap with other granulocytes and PBMCs. Results: The machine learning tool achieved 99.85% sensitivity and 99.81% specificity in neutrophil classification on public test sets, navigating low RNA and fragility challenges. By integrating marker-based and machine learning annotations, our pipeline pinpointed robust neutrophil signatures in our dataset, priming biological insights. Discussion: This work demonstrates the power of ML in resolving cellular complexity in liquid biopsies and lays the groundwork for future development of neutrophil-based biomarkers in TNBC. Additionally, it addresses a critical gap in current methodologies by offering a robust, reproducible strategy to overcome the inherent challenges in neutrophil identification and classification from peripheral blood. Citation Format: Krzysztof Pastuszak, Marcin Banacki, Michał Sieczczyński, Anna Supernat, Anna J. Żaczek. Machine learning used to validate neutrophil classification in Triple Negative Breast Cancer patients [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Artificial Intelligence and Machine Learning; 2025 Jul 10-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(13_Suppl):Abstract nr B038.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".