Abstract B027: The development of a multiscale transcriptional atlas of sarcoma
Notice bibliographique
Résumé
Abstract Objectives: Sarcomas are mesodermal cancers of bone and soft tissue of which there are >60 malignant varieties, many of which can be difficult to diagnose or subtype using traditional histopathology. A universal molecular definition of sarcoma types would therefore be an invaluable tool to the diagnostic pathologist. RNA has the potential to offer a complimentary perspective to cytogenetic- and methylation-based diagnostics, as it represents the active state of the disease at sampling and better reveals its phenotype. Recognizing the potential for RNA-based classification, we set out to create a first-generation transcriptional atlas of sarcoma. Methods: To develop transcriptional definitions of cancers with the potential to further subclassify tumor types, we designed a self-optimizing and scale-adaptive unsupervised method (RACCOON), which groups samples into hierarchically organized clusters. We used this approach on the UCSC Treehouse Childhood Cancer Compendium, a set of 2,178 pediatric and 9,400 adult tumors, 1,130 of which are sarcomas, as well as 1,735 non-neoplastic samples. We then trained an ensemble of convolutional neural networks to classify tumors to these transcriptional clusters. We have now added 624 more sarcoma samples from Toronto centers and international collaborators to better represent the breadth of sarcoma. We are actively sequencing 500 additional samples in partnership with the Gabriella Miller Kids First Research Program to yield an expanded cohort of >2,200 uniformly processed and analyzed sarcomas. Results: Sarcomas organize into two clusters at the highest hierarchical level: one characterized by entities which occur primarily in adults and resemble mature tissue, the other by primarily pediatric entities which exhibit high stemness and resemble embryonic tissue. Several included entities are not bona fide sarcomas but originate from the mesoderm (e.g., Wilms Tumor) signifying a common transcriptional identity for mesodermal neoplasms. Additionally, we demonstrate the first transcriptional subtypes of central osteosarcoma reflecting its major histotypes and representing divergent clinical courses. We also determine Ewing Sarcoma (ES) to be a distinct entity which clusters separately from all other cancers, raising questions of its origin and affinity to sarcoma. When classifying ongoing patients to the atlas, we correctly classified >85% of tumors and corrected the diagnosis of 7%. We find 14% of ES in our dataset were likely misdiagnosed CIC- or BCOR-driven sarcomas. Critically, assigned subtypes are consistent between primary and relapse pairs. Conclusion: RNA-seq is a promising tool for both subtype discovery and classifying sarcoma in ongoing patients. We have already included this tool in tumor boards to help inform patient care. Our method reveals the overarching organization of sarcoma for the first time and specifies its underlying biology. This atlas is ever-growing and is open to the community to contribute. Citation Format: Joshua O. Nash, Federico Comitani, Rose Chami, Sarah Cohen-Gogo, Astra Chang-Schwertschkow, Yael Babichev, Jodi Lees, Noa Alon, Nalan Gokgoz, Stephen Man Yu, Kyoko Yuki, Miranda Lorenti, Zhanqin Liu, Alaina McGoey, Famida Spatare, Bernarld Castro, Kim Tsoi, Hagit Peretz Soroka, Jack Brzezinski, Anita Villani, Albiruni Razak, Abha Gupta, Elizabeth Demicco, Gino Somers, Brendan C. Dickson, Jay S. Wunder, Irene L. Andrulis, David Malkin, Rebecca A. Gladdy, Adam Shlien. The development of a multiscale transcriptional atlas of sarcoma [abstract]. In: Proceedings of the AACR Special Conference: Sarcomas; 2022 May 9-12; Montreal, QC, Canada. Philadelphia (PA): AACR; Clin Cancer Res 2022;28(18_Suppl):Abstract nr B027.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».