Abstract A016: ews-nf: A custom workflow for tumor cell annotation and analysis of single-cell RNA-sequencing of paired patient Ewing sarcoma specimens from the Sean Karl cohort
Notice bibliographique
Résumé
Abstract Ewing sarcoma (EwS) is a fusion oncoprotein-driven primary bone cancer that demonstrates vast intra- and inter-tumoral heterogeneity. Tumor cell subpopulations, tumor progression, and therapeutic vulnerabilities of cell subpopulations are poorly understood. Paired patient samples (primary and metastatic disease or relapse) are rare at any one institution, and collaborative efforts are needed to address these pressing biologic questions. A national collaborative effort (the Sean Karl cohort) has been established to conduct single-cell RNAseq analyses of retrospective paired tumor samples from patients with EwS and serial samples due to metastasis or relapse using the GEM-X Flex Gene Expression protocol from 10x Genomics. Four analytic teams from multiple institutions will be performing custom downstream analyses to understand the therapeutic vulnerabilities of EwS cell subsets and discern immunobiologic dysfunction. To ensure reproducibility, all data pre-processing and common analyses, such as cell type annotation, are centralized using reproducible workflows developed by the Childhood Cancer Data Lab, a program of Alex’s Lemonade Stand Foundation. Gene expression is quantified using an open-source workflow, scpca-nf. The output from scpca-nf, which includes raw and normalized gene expression, dimensionality reduction, and annotation of non-malignant cells, is used as input to a custom Nextflow workflow, ews-nf, to annotate and analyze tumor cells in Ewing sarcoma samples. Tumor cells are annotated using two complementary methods: AUCell is used to evaluate expression of EwS-specific gene sets, and inferCNV is used to obtain a CNV profile for each cell by comparing potentially malignant cells to definitively non-malignant cells (e.g., immune cell types). Cells with high expression of EwS-specific gene sets and high CNV profiles, relative to immune cell types, are annotated as tumor cells. Tumor cells are further divided into EWS::FLI1 “low” and “high” cells based on expression of custom gene sets. All tumor cells are then analyzed using non-negative matrix factorization to identify and label recurrent gene expression programs found across all samples in the cohort. Processing and sequencing of samples are ongoing at the time of abstract submission. Ultimately, the output from ews-nf will be used to create a harmonized dataset to be shared with all four analytic teams. This harmonized dataset will contain the processed gene expression data, labeling of tumor cells and tumor cell states, and identification of recurrent gene expression programs. This enables all analytical teams to conduct downstream analysis using the same set of tumor cell annotations, making it easy for teams to compare results and draw conclusions. After completion of the study, the ews-nf workflow will be made publicly available to the research community. The processed gene expression data from the Sean Karl cohort, including the tumor cell annotations, will also be made available on the Single-cell Pediatric Cancer Atlas Portal for others to use in their own research. Citation Format: Allegra G Hawkins, Stephanie J Spielman, Joshua A Shapiro, Abbe Pannucci, Elina Mukherjee, Jessica Daley, Shireen Ganapathi, Elissa Boguslawski, Lea F Surrey, Patrick Azar, Filemon Dela Cruz, Jovana Pavisic, Emily Stockfisch, Azfar Neyaz, Ivy John, Jennifer Picarsic, Yutaro Tanaka, Riaz Gillani, Katherine A Janeway, Jaclyn N Taroni, Jessica Davis, Damon Reed, Adam Shlien, Theodore Laetsch, Rajen Mody, Elizabeth R Lawlor, Patrick Grohar, Anthony R Cillo, Kelly M Bailey. ews-nf: A custom workflow for tumor cell annotation and analysis of single-cell RNA-sequencing of paired patient Ewing sarcoma specimens from the Sean Karl cohort [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Discovery and Innovation in Pediatric Cancer— From Biology to Breakthrough Therapies; 2025 Sep 25-28; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(18_Suppl_2):Abstract nr A016.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,005 | 0,005 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,002 |
| Bibliométrie | 0,002 | 0,001 |
| Études des sciences et des technologies | 0,002 | 0,001 |
| Communication savante | 0,003 | 0,001 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,048 | 0,032 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».