Abstract PR16: Comprehensive transcriptomic characterization of 1,400 sarcomas for diagnosis and immune contexture
Bibliographic record
Abstract
Abstract Introduction: Sarcomas represent about 15 % of all childhood cancers and are still lethal in a large proportion of cases. They comprise a heterogeneous group of bone and soft tissue tumors with more than 100 distinct histologic entities, often presenting a diagnostic challenge to pathologists. Methods: To address this pathological complexity and better characterize the molecular biology of sarcomas, we performed whole-transcriptome RNA-seq on frozen tumor tissues as part of the diagnostic workflow for more than 1,400 patients, mainly children and young adults, addressed for sarcoma molecular diagnosis throughout France. Results: From a clinical perspective, RNA-seq proves to be a powerful complementary test to pathology for diagnosis of sarcomas: it allows agnostic detection of subtype-specific fusion transcripts, as well as diagnostic indication by gene-expression profile clustering. In addition to its diagnostic utility, RNA-seq provides a rich source of information for studying the molecular biology of sarcomas. Indeed, as shown recently in our group, it allows detection of previously uncharacterized sarcoma entities associated with novel fusion genes. Using dimensionality reduction techniques, either linear such as PCA, or nonlinear such as t-SNE or UMAP, we perform unsupervised clustering of our cohort, revealing insights into the global structure of the transcriptomic landscape of sarcomas. We show that many histologic entities have a distinct gene-expression profile, emphasizing the utility of RNA-seq for diagnosis. To explore the immune microenvironment of sarcomas and potential correlates of response to immunotherapy, we deconvolute bulk RNA-seq data to estimate the presence of immune and stromal cell types, allowing a comprehensive characterization of the immune landscape of sarcomas. Finally, we apply state-of-the-art deep learning tools to perform unsupervised clustering and nonlinear dimensional reduction of our RNA-seq samples, allowing a finer and more insightful representation of the structure of the sarcoma transcriptomic landscape. Conclusion: Besides improving the diagnostic workflow for patients, our comprehensive analysis of sarcoma RNA-seq allows characterization of distinct molecular subtypes and of the immune landscape, opening up perspectives for better diagnosis, classification, and biologic discoveries in sarcoma. This abstract is also being presented as Poster B75. Citation Format: Julien Vibert, Sarah Watson, Camille Benoist, Joshua Waterfall, Gaëlle Pierron, Olivier Delattre. Comprehensive transcriptomic characterization of 1,400 sarcomas for diagnosis and immune contexture [abstract]. In: Proceedings of the AACR Special Conference on the Advances in Pediatric Cancer Research; 2019 Sep 17-20; Montreal, QC, Canada. Philadelphia (PA): AACR; Cancer Res 2020;80(14 Suppl):Abstract nr PR16.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".