MétaCan
Menu
← Back to cohort
Record W4413377061 · doi:10.1101/2025.08.12.25333156

SuReCAN: a suite of user-friendly Galaxy machine learning workflows to predict survival and treatment response of cancer patients

2025· preprint· en· W4413377061 on OpenAlexfundno aff
Jie Ju, Daria Koppes, Andrew Stubbs, Yunlei Li

Bibliographic record

VenuemedRxiv · 2025
Typepreprint
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsnot available
FundersGovernment of Ontario
KeywordsSuiteWorkflowGalaxyComputer scienceUser FriendlyHuman–computer interactionAstrophysicsOperating systemPhysicsDatabaseGeography

Abstract

fetched live from OpenAlex

Abstract Cancer is one of the leading lethal causes worldwide, with enormous impact on healthcare, economy and society. One of the main challenges of clinical treatment planning is that patients usually have diverse clinical outcomes given the same diagnosis and treatments. To enable personalized cancer therapeutic planning, (bio)medical data analyses using machine learning (ML) models are introduced to efficiently extract informative biological patterns from the massive volume of complex biological data, aiding in cancer patients’ stratifications. For biomedical researchers without computational biology background, the gap between clinical practice and computational approaches is prominent and hinders the usage of machine learning in medical research. To fill this gap, we created a collection of ML workflows on the Galaxy platform named SUrvival and REsponse prediction for CANcer patients (SuReCAN) for clinicians and biologists to build and deploy predictive ML classifiers. Being freely available and accessible, SuReCAN automates the data analysis process and enables the clinicians and researchers to perform a broad range of predictive tasks. It contains a toolkit of three ML modules with various existing and newly implemented methods on Galaxy: A data normalization module, a feature selection module, and an ML classifier module. We exhibited the utility of SuReCAN with a few real-world datasets to identify pancreatic ductal adenocarcinoma (PDAC) patients’ survival-correlated subtypes and to predict drug response outcomes based on various omics data from patient tumor samples. As a result, all workflows achieved a median accuracy of over 0.8 in PDAC survival-correlated subtype classification. In particular, the workflow combining the feature selection method “SVM-based RFECV ” and the Support Vector Machine classifier consistently outperformed the other workflows, while all classifiers have shown their superiority on different omics data. Importantly, SuReCAN is not only applicable for the clinical prediction tasks shown in the test cases but also suitable for new classifier development and deployment with clinical observations provided by the users. Providing a collection of user-friendly ML workflows, SuReCAN stratifies patients based on their biomedical profiling in a data-driven way and assists biomedical researchers with clinical decision-making and scientific discoveries.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.008
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.012
Threshold uncertainty score0.029

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0030.008
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.002
Bibliometrics0.0020.001
Science and technology studies0.0010.001
Scholarly communication0.0020.001
Open science0.0020.002
Research integrity0.0010.001
Insufficient payload (model declined to judge)0.0090.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.013
GPT teacher head0.303
Teacher spread0.290 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venuemedRxiv→Same topicRadiomics and Machine Learning in Medical Imaging→French-language works237,207→