Abstract 2500: An automated and scalable pipeline for high-dimensional immune cell phenotyping in mass cytometry datasets
Bibliographic record
Abstract
Context: With the increasing use of spectral flow and mass cytometry technologies, efficient single-cell phenotyping has become essential for the identification of complex cell populations. Traditionally, these populations are identified through manual gating, a time-consuming process that is also subject to variability and a lack of reproducibility. Here, we adapted a deep learning model [1] to devise an automated gating pipeline with the goal of enhancing the accuracy, speed, and reproducibility of immune cell gating. Methods: Our pipeline was evaluated on two mass-cytometry datasets [1], [2] (>5M cells, 35 markers), that had been previously gated by experts for comparison purposes. Phenotypic markers were used for identification, while the median expression of intracellular markers was used to determine the functional properties of populations. Accuracy and F1-score were used to compare classified populations against manually gated populations. A third dataset (15M events from 57 patients diagnosed with lung cancer) was used to compare the prediction of lung cancer progression after immunotherapy using properties based on manually vs. automatically gated populations. Results: The automated gating pipeline achieved high overall accuracy comparable to expert manual gating (n=30 cell populations) (0.917 and 0.921 for dataset 1 and 2, respectively). Processing time was significantly reduced: <12min for both datasets (8CPUs, 8GB of RAM). High-level populations (n=10) were identified with excellent accuracy (e.g. Tcells, Bcells with f1-score of 0.993, 0.937 on the first dataset and 0.996, 0.934 on the second dataset). However, rare populations (n=20) showed higher discrepancies (e.g. intermediate monocytes, DC with f1-scores of 0.864, 0.816 and 0.678, 0.552 on dataset 1 and 2, respectively). Differences in identification scores had low effect on the functional properties of populations, as 91.4% of the functional properties defined from our pipeline were highly correlated (r>0.75) with those derived from manual gating. Further, the prediction of lung cancer progression after immunotherapy showed similar or improved results using functional properties based on automated identified populations (AUC=0.82) compared to manually gated populations (AUC=0.70). Conclusion: This automated approach eliminates operator biases and handles multiple markers simultaneously, offering reliable and efficient analyses for research and clinical applications. To explore discrepancies, future work will incorporate datasets annotated by multiple experts to assess inter-expert variability. 1. Blampey, Q. et al, A biology-driven deep generative model for cell-type annotation in cytometry. Brief Bioinform. 2023 Sep 20;. doi: 10.1093/bib/bbad260. 2. https://clinicaltrials.gov/study/NCT05523713 3. Ina A. Stelzer et al, Sci.Transl.Med.13, (2021). DOI:10.1126/scitranslmed.abd9898 Citation Format: Benjamin Waked, Grégoire Bellan, Xavier Durand, Alexandre Maillard, Franck Verdonk, Brice Gaudilliere, Helen McGuire, Natalie Smith, Christina Loh, David King, Dominique Blanchard, Julien Hedou. An automated and scalable pipeline for high-dimensional immune cell phenotyping in mass cytometry datasets [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 2500.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.003 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.008 | 0.009 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".