MétaCan
Menu
Retour à la cohorte
Enregistrement W4282939738 · doi:10.1158/1538-7445.am2022-3813

Abstract 3813: Unification of over 50 published single-cell RNA datasets covering over a 1000 patient samples with deep-learning reveals novel axes of tumor microenvironment variation

2022· article· en· W4282939738 sur OpenAlexaff
Javier Díaz-Mejía, X Swechha, Dylan Mendonca, Octavian Focsa, Chris J. Harvey, Mike Briskin, Sam Cooper

Notice bibliographique

RevueCancer Research · 2022
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueSingle-cell and spatial transcriptomics
Établissements canadiensToronto Centre for Phenogenomics
Organismes subventionnairesnon disponible
Mots-clésComputer scienceWorkflowGround truthArtificial intelligenceAtlas (anatomy)Deep learningMachine learningComputational biologyBiologyDatabase

Résumé

récupéré en direct d'OpenAlex

Abstract Background: Due to sizeable batch effect challenges, published scRNA datasets remain siloed, with no tools or packages yet demonstrating an ability to integrate more than a handful of datasets into a unified atlas. This has precluded generation of a large-unified Atlas akin to The Cancer Genome Atlas (TCGA) for scRNA, despite massive demand for such a resource. Methods: We built the first training scRNA dataset that we know of that relies on using cell-type labels from published studies as a ground-truth metric. Using this dataset, we evaluated and trained a variety of models specifically for the task of integrating disparate data into a unified space. Further we developed a specific framework for evaluating how well unsupervised models perform at the task of integrating disparate data, using a new approach reliant on leave-one out validation of ‘unseen’ datasets. Using deep-learning models, that performed best on the training dataset, we scaled integration of over 50 public datasets focused on solid cancers, that collectively contain over 1000 patient samples worth of data covering over 20 indications. Results: The pan-cancer scRNA atlas produced by the above workflow is an order of magnitude larger than previous scRNA datasets and the first to span many indications alongside adjacent and separate normal tissue data. Analysis of this atlas reveals novel axes of variation in the tumor microenvironment linked to Cancer Associated Fibroblast (CAF) biology. For example: (a) We find CAF high samples vs. cancer high samples are enriched for T-cells in a naïve state; (b) Cancer vs. CAF rich samples result in variation in M2 like macrophage signatures; this compartmentalization is also seen in spatial RNA data; (c) A spectrum of CAF, perivascular, and endothelial like states is also observed indicating potential cell-type plasticity. Collectively, these observations identify novel biology and variation in the tumor microenvironment that will likely apply to many ongoing experimental projects and therapeutic programs. Conclusions: We’ve used deep-learning to build one of the largest scRNA atlases to date, and potentially the first that will progressively release models and data as an open-source package to benefit the wider community. We anticipate the ability to resolve target expression at the single cell level will greatly enhance our understanding of the tumor microenvironment, as it aids our own efforts to drug CAF biology and the tumor stroma. Citation Format: Javier Díaz-Mejía, Swechha X, Dylan Mendonca, Octavian Focsa, Chris Harvey, Mike Briskin, Sam Cooper. Unification of over 50 published single-cell RNA datasets covering over a 1000 patient samples with deep-learning reveals novel axes of tumor microenvironment variation [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 3813.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,001
score de la tête « metaresearch » (Gemma)0,000
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Expérimental (laboratoire) · Signal consensuel: Expérimental (laboratoire)
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,050
Score d'incertitude au seuil0,513

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0010,000
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0000,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0000,000
Intégrité de la recherche0,0000,000
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,033
Tête enseignante GPT0,278
Écart entre enseignants0,245 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeExpérimental (laboratoire)
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueCancer ResearchMême sujetSingle-cell and spatial transcriptomicsTravaux en français237 207