MétaCan
Menu
Retour à la cohorte
Enregistrement W4282939738 · doi:10.1158/1538-7445.am2022-3813

Abstract 3813: Unification of over 50 published single-cell RNA datasets covering over a 1000 patient samples with deep-learning reveals novel axes of tumor microenvironment variation

2022· article· en· W4282939738 sur OpenAlexaff
Javier Díaz-Mejía, X Swechha, Dylan Mendonca, Octavian Focsa, Chris J. Harvey, Mike Briskin, Sam Cooper

Notice bibliographique

RevueCancer Research · 2022
Typearticle
Langueen
DomaineBiochemistry, Genetics and Molecular Biology
ThématiqueSingle-cell and spatial transcriptomics
Établissements canadiensToronto Centre for Phenogenomics
Organismes subventionnairesnon disponible
Mots-clésComputer scienceWorkflowGround truthArtificial intelligenceAtlas (anatomy)Deep learningMachine learningComputational biologyBiologyDatabase

Résumé

récupéré en direct d'OpenAlex

Abstract Background: Due to sizeable batch effect challenges, published scRNA datasets remain siloed, with no tools or packages yet demonstrating an ability to integrate more than a handful of datasets into a unified atlas. This has precluded generation of a large-unified Atlas akin to The Cancer Genome Atlas (TCGA) for scRNA, despite massive demand for such a resource. Methods: We built the first training scRNA dataset that we know of that relies on using cell-type labels from published studies as a ground-truth metric. Using this dataset, we evaluated and trained a variety of models specifically for the task of integrating disparate data into a unified space. Further we developed a specific framework for evaluating how well unsupervised models perform at the task of integrating disparate data, using a new approach reliant on leave-one out validation of ‘unseen’ datasets. Using deep-learning models, that performed best on the training dataset, we scaled integration of over 50 public datasets focused on solid cancers, that collectively contain over 1000 patient samples worth of data covering over 20 indications. Results: The pan-cancer scRNA atlas produced by the above workflow is an order of magnitude larger than previous scRNA datasets and the first to span many indications alongside adjacent and separate normal tissue data. Analysis of this atlas reveals novel axes of variation in the tumor microenvironment linked to Cancer Associated Fibroblast (CAF) biology. For example: (a) We find CAF high samples vs. cancer high samples are enriched for T-cells in a naïve state; (b) Cancer vs. CAF rich samples result in variation in M2 like macrophage signatures; this compartmentalization is also seen in spatial RNA data; (c) A spectrum of CAF, perivascular, and endothelial like states is also observed indicating potential cell-type plasticity. Collectively, these observations identify novel biology and variation in the tumor microenvironment that will likely apply to many ongoing experimental projects and therapeutic programs. Conclusions: We’ve used deep-learning to build one of the largest scRNA atlases to date, and potentially the first that will progressively release models and data as an open-source package to benefit the wider community. We anticipate the ability to resolve target expression at the single cell level will greatly enhance our understanding of the tumor microenvironment, as it aids our own efforts to drug CAF biology and the tumor stroma. Citation Format: Javier Díaz-Mejía, Swechha X, Dylan Mendonca, Octavian Focsa, Chris Harvey, Mike Briskin, Sam Cooper. Unification of over 50 published single-cell RNA datasets covering over a 1000 patient samples with deep-learning reveals novel axes of tumor microenvironment variation [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 3813.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,003
score de la tête « metaresearch » (Gemma)0,005
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: aucune
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,005
Score d'incertitude au seuil0,015

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0030,005
Méta-épidémiologie (sens strict)0,0010,000
Méta-épidémiologie (sens large)0,0010,001
Bibliométrie0,0020,001
Études des sciences et des technologies0,0010,001
Communication savante0,0010,001
Science ouverte0,0010,002
Intégrité de la recherche0,0010,001
Charge utile insuffisante (le modèle a refusé de juger)0,0020,001

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,033
Tête enseignante GPT0,278
Écart entre enseignants0,245 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2022
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueCancer ResearchMême sujetSingle-cell and spatial transcriptomicsTravaux en français237 207