MétaCan
Menu
Back to cohort
Record W4282983957 · doi:10.1158/1538-7445.am2022-1221

Abstract 1221: Comprehensive cell-type classification of tumor and normal cells from single cell RNA sequencing in pan cancer settings

2022· article· en· W4282983957 on OpenAlexaff
Ido Nofech-Mozes, Philip Awadalla, Sagi Abelson

Bibliographic record

VenueCancer Research · 2022
Typearticle
Languageen
FieldBiochemistry, Genetics and Molecular Biology
TopicSingle-cell and spatial transcriptomics
Canadian institutionsOntario Institute for Cancer Research
Fundersnot available
KeywordsCell typeTranscriptomeAnnotationClassifier (UML)Computational biologyCancer cellCancerBiologyCellStromal cellGeneBioinformaticsGene expressionArtificial intelligenceCancer researchComputer scienceGenetics

Abstract

fetched live from OpenAlex

Abstract Single-cell RNA sequencing (scRNA-seq) allows for the study of the transcriptome at a cellular level, where populations of cells are annotated based on the expression of marker genes, providing a tool to gain cell-specific, subtle insights on cancer biology. However, precise annotation of cell type remains a challenge, hindering the efficiency of data interpretation. Several existing tools for cell-type annotation have been developed to improve resolution and reproducibility, yet their performance is reduced when the reference dataset contains many cell types, subclasses of similar cell types, or malignant cells. Interpatient malignant cell heterogeneity often leads to reduced accuracy when classifying cancer cells as most methods rely on correlations to a reference from a different source. Given the challenges in the annotation of scRNA-seq data of cancer and its high impact for elucidating mechanisms associated with tumor heterogeneity, pathogenesis, and treatment, we developed a comprehensive, hierarchically organized, multi-layered classifier spanning diverse malignant and normal cells of the tumor microenvironment. We found that performance improves when each layer focuses on a smaller number of classes and each cell sequentially moves down a series of classifiers with increased cell type resolution. When applied to an external validation dataset of over 300 primary solid tumor biopsies spanning diverse cancer types, the classifier accurately annotated the tissue of origin of malignant cells, and relevant subtypes of stromal and blood cells, with average F1 scores of 0.91, 0.95 and 0.99 respectively. Using confidence thresholds at each layer, the classifier abstains from classifying ambiguous cells. We applied 4 existing annotators provided with the same reference to the external test dataset and found that cancer cells are misclassified or unclassified, while the blood and stromal cells are accurately classified, highlighting our tool’s unique ability to classify cancer cells. Moreover, we applied our classifier’s to scRNA-seq data derived from breast cancer metastasis to the liver and were able to uncover the tissue of origin, demonstrating the potential use for determining the source of a metastatic tumour. Finally, given that our classifier is modular, we leveraged two recently published single cell breast cancer atlases to add a breast cancer subtype classification layer, that consistently identified the correct clinical subtype of single breast cancer cells in external data. This study provides a flexible model for the annotation of cells comprising the tumor microenvironment in pan cancer settings, while existing methods require tissue-specific references for every cancer type. Our classifier provides a powerful method for investigating intercellular communication pathways between tumor cells and non-malignant cells of the tumor microenvironment. Citation Format: Ido Nofech-Mozes, Philip Awadalla, Sagi Abelson. Comprehensive cell-type classification of tumor and normal cells from single cell RNA sequencing in pan cancer settings [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1221.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: Bench or experimental
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.007
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.073
GPT teacher head0.322
Teacher spread0.249 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueCancer ResearchSame topicSingle-cell and spatial transcriptomicsFrench-language works237,207