MétaCan
Menu
← Back to cohort
Record W4411917217 · doi:10.1101/2025.06.30.25330543

Advancing sarcoma diagnostics with expanded DNA methylation-based classification

2025· preprint· en· W4411917217 on OpenAlexaff
Natalie Jäger, David Reuß, Martin Sill, Daniel Schrimpf, Abigail K. Suwala, Philipp Sievers, Rouzbeh Banan, Felix Hinz, Ramin Rahmanzade, Henry Bogumil, Kaan Fuat Aras, Areeba Patel, Andrey Korshunov, Melanie Bewerunge‐Hudler, Arjen H.G. Cleven, Manel Esteller, Hanno Glimm, Wolfgang Hartmann, Simon Kreutzfeld, Christoph E. Heilig, Till Milde, Iver Petersen, Wolfgang Wick, Olaf Witt, Thibault Kervarrec, Evelina Miele, Jonathan Serrano, Stephan Frank, Karl Kashofer, Anne Mc Leer, Elke Pfaff, Mélanie Pagès, Arnault Tauziède‐Espariat, Ferdinand Toberer, Henning B. Boldt, Petr Martínek, Sebastian Brandner, Mayara Ferreira Euzébio, Aurore Siegfried, Jane Chalker, P Harter, Romain Appay, Wolfgang Dietmaier, Martin Hasselblatt, Uta Flucke, Laura S. Hiemcke‐Jiwa, David A. Solomon, Clara Frydrychowicz, Pascale Varlet, Benjamin Goeppert, Michaela Nathrath, Claudia Blattmann, Monika Sparber‐Sauer, Michel Mittelbronn, Thomas Mentzel, Sandra Leisz, Anja Harder, Till Acker, Drew Pratt, Eva Wardelmann, Jamal Benhamida, M. Ladanyi, Philipp Jurmeister, William D. Foulkes, Pamela Ajuyah, Jürgen Hench, Maikel JL. Nederkoorn, Yvonne M.H. Versleijen‐Jonkers, Gunhild Mechtersheimer, Sandro M. Krieg, Manfred Gessler, Daniel Baumhoer, Sam Behjati, Luca Bertero, Klaus Griwank, Dirk Schadendorf, Pancras C.W. Hogendoorn, Jean‐François Emile, Paul G. Kemps, Armin Jarosch, Michael Ronellenfitsch, Toni Su Idler, Daniela E. Aust, Sylvia Herold, Jessica Pablik, Maysa Al‐Hussaini, Zied Abdullaev, Maximus C.F. Yeung, Marco Wachtel, Eva Brack, F. Kommoss, Markku Miettinen, Ken Aldape, Adrienne M. Flanagan, Uta Dirksen, Kristian W. Pajtler, Thomas G. P. Grünewald, Daniel B. Lipka, Stefan Fröhling, Christian Koelsche, Matija Snuderl, David Capper, Stefan M. Pfister, David Jones, Felix Sahm, Andreas von Deimling

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldMedicine
TopicSarcoma Diagnosis and Treatment
Canadian institutionsMcGill University
FundersDeutsche KinderkrebsstiftungDeutsche KrebshilfeDeutschen Konsortium für Translationale KrebsforschungBundesministerium für GesundheitBundesministerium für Bildung und ForschungDeutsches Krebsforschungszentrum
KeywordsSarcomaClassifier (UML)Artificial intelligenceMachine learningPathologyOncologyMedicineComputer science

Abstract

fetched live from OpenAlex

Purpose: Sarcomas pose a severe diagnostic challenge. A wide variety of these distinct entities need to be distinguished from each other and from less aggressive types of mesenchymal tumors, to ensure correct clinical management. A machine learning based classifier for sarcomas utilizing DNA methylation data from 1077 tumors recognizing 62 sarcoma types has already been developed and termed the sarcoma classifier, which we published in 2021. Here we present a major advancement of the scale and precision of the sarcoma classifier. Methods: DNA methylation profiles and histologic data from an unprecedented multi-institutional cohort of mesenchymal tumors were collected and analyzed. Utilizing a machine learning approach, the classifier was rigorously validated through five-fold nested cross-validation, achieving a 98% class-level accuracy and a Brier score of 0.017, indicative of well-calibrated probability estimates. Results: The sarcoma classifier v13.1 was developed based on a training set of 4377 methylation profiles from sarcomas and less aggressive mesenchymal tumors comprising 116 tumor sub-classes and 4 control groups forming 93 distinct methylation classes. Performance was validated using four independent cohorts, comprising a total of 1547 mesenchymal tumors. A methylation-based classifier prediction was obtained in 73% of cases in the validation sets, of which 91% matched the original histopathology diagnosis, thereby increasing diagnostic confidence. The classifier enabled a definitive molecular diagnosis or tumor reclassification in 6% of cases with inconclusive or ambiguous histological findings. Conclusion: Adding new sarcoma types and expanding tumor sample numbers in each methylation class in the new sarcoma classifier decisively increased the number of diagnostic predictions and improved match with histologic evaluation. This substantial advancement will promote clinical implementation of the tool for the diagnosis of mesenchymal tumor lesions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.007
metaresearch head score (Gemma)0.015
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.007
Threshold uncertainty score0.036

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0070.015
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.001
Science and technology studies0.0000.001
Scholarly communication0.0010.001
Open science0.0010.001
Research integrity0.0010.002
Insufficient payload (model declined to judge)0.0020.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.030
GPT teacher head0.310
Teacher spread0.279 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuemedRxiv→Same topicSarcoma Diagnosis and Treatment→French-language works237,207→