MétaCan
Menu
Back to cohort

The use of machine learning models to predict PFS and OS outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES).

2023· article· en· W4385547547 on OpenAlexaff
Khadjah Alshankati, Aisha Alshibany, Augustin Toma, Katherine Lajkosz, Bo Wang, Benjamin Haibe‐Kains, Lillian L. Siu

Bibliographic record

VenueJCO Global Oncology · 2023
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversity Health NetworkPrincess Margaret Cancer CentreUniversity of Toronto
Fundersnot available
KeywordsWaterfallMedicineClinical endpointRandomized controlled trialSample size determinationClinical trialStatisticsArtificial intelligenceInternal medicineComputer scienceMathematics

Abstract

fetched live from OpenAlex

107 Background: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is an emerging short-term endpoint that may represent a surrogate for survival-based outcomes such as PFS and OS. We hypothesize that the configuration of waterfall plots in randomized trials may predict PFS/OS outcomes. Methods: A literature-based search was performed for all phase II/III randomized clinical trials published in MEDLINE from 2010 to 2022 testing molecularly targeted agents (MTA) or immunotherapy (IO). Articles reporting at least 1 waterfall plot for each treatment arm depicting maximum DepOR of target lesions with corresponding PFS/OS Kaplan-Meier plots were included. Studies are defined as positive or negative based on the achievement of a priori stated primary endpoint. Trial data collected included sample size per arm, cancer type, mechanisms of action of drug(s) tested, line of treatment, etc. Images of waterfall plots were manually extracted from publications and then processed through a semi-automatic extraction process using WebPlotDigitizer and Tesseract to produce tabular representations. Logistic regression with L2 regularization was used for modeling; hyperparameter tuning was accomplished with five-fold cross-validation on a training set compromising 80% of the data. Results: A total of 111 studies were identified: 65 (59%) phase III and 46 (41%) phase II, mean sample size per arm 317 (19-1581). Most frequent cancer type was gastrointestinal 24 (22%). MTA, IO and combinations were tested in 113 (51%), 35(16%) and 11 (5%) studies respectively. Chemotherapy and other treatment regimens were used in 63 (28%) trials. PFS was the primary endpoint in 62 (56%); 80 (75%) studies were positive. Of the 111 studies only 83 (75%) were retained for machine learning analysis, the remainder were excluded due to atypical formatting such as superimposed waterfall plots. Performance of the model was assessed on a test set which comprised 20% of the original dataset. Table below shows the classification metrics from modelling. Both PPV and NPV were ≥80%. Conclusions: MAP-OUTCOMES evaluated pan-cancer randomized studies with diverse therapeutic anticancer agents. It is a computational tool with the potential to predict survival-based outcomes from waterfall plots and may help with decisions regarding follow-on randomized studies. Further validation is ongoing. [Table: see text]

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.014
metaresearch head score (Gemma)0.035
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.609
Threshold uncertainty score0.973

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0140.035
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0030.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.180
GPT teacher head0.445
Teacher spread0.265 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJCO Global OncologySame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207