MétaCan
Menu
Back to cohort
Record W4412440980 · doi:10.1016/j.esmoop.2025.105509

The use of machine learning models to predict progression-free survival and overall survival outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES)

2025· article· en· W4412440980 on OpenAlexaff
Khadjah Alshankati, Aisha Alshibany, Azhar Toma, Katherine Lajkosz, Benjamin Haibe‐Kains, Lillian L. Siu

Bibliographic record

VenueESMO Open · 2025
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsVector InstitutePrincess Margaret Cancer CentreUniversity Health Network
Fundersnot available
KeywordsWaterfallClinical endpointConfidence intervalMedicineRandomized controlled trialProgression-free survivalLogistic regressionWaterfall modelArtificial intelligenceMachine learningInternal medicineStatisticsOncologyOverall survivalComputer scienceMathematicsSoftwareCartography

Abstract

fetched live from OpenAlex

BACKGROUND: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is a short-term endpoint that may represent a surrogate for survival-based outcomes such as progression-free survival (PFS) and overall survival (OS). We hypothesized that PFS/OS could be predicted from waterfall plots in randomized clinical trials (RCTs) using a novel machine-learning (ML) computational model. MATERIALS AND METHODS: A literature-based search was carried out for phase II/III RCTs testing noncytotoxic systemic therapy, which included waterfall plots with corresponding PFS/OS results. Studies were defined as positive or negative based on achievement of an a priori-stated primary endpoint. Trial data and images of waterfall plots were manually extracted and then processed through a semi-automatic extraction process. We developed the MAP-OUTCOMES (MAchine learning model to Predict PFS and OS OUTCOMES) model using regularized logistic regression. This model was applied to a training set comprising 70% of the data, and 30% was used for a test set. RESULTS: A total of 91 unique RCTs were identified, and 82 (93 trial pairs) retained for the ML analysis. Most of the trials were phase III (75%), with 67% using PFS as the primary endpoint and a mean sample size of 350 patients per arm. The most common tumor type was genitourinary (22%), and small-molecule targeted agents (27%) were the most frequent regimen. The model's performance achieved 71% accuracy [95% confidence interval (CI) 0.536-0.862, P = 0.18] with an area under the curve (AUC) of 65% (95% CI 0.333-0.938, P = 0.157) and area under the precision-recall curve (AUPRC) of 90% (95% CI 0.779-0.995, P = 0.171) in the 28 trials used for the test set. CONCLUSIONS: The MAP-OUTCOMES model demonstrated the feasibility of using ML to predict survival-based outcomes from waterfall plots, thus providing a potential tool for early trial evaluation. Improving the model's performance with more training data and creating independent datasets are necessary steps to assess its generalizability for prospective clinical applications.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.025
metaresearch head score (Gemma)0.046
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: Observational
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.514
Threshold uncertainty score0.996

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0250.046
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0030.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0010.001
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.147
GPT teacher head0.428
Teacher spread0.281 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueESMO OpenSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207