MétaCan
Menu
Back to cohort

The use of machine learning models to predict PFS and OS outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES).

2023· article· en· W4385547547 on OpenAlexaff
Khadjah Alshankati, Aisha Alshibany, Augustin Toma, Katherine Lajkosz, Bo Wang, Benjamin Haibe‐Kains, Lillian L. Siu

Bibliographic record

VenueJCO Global Oncology · 2023
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversity Health NetworkPrincess Margaret Cancer CentreUniversity of Toronto
Fundersnot available
KeywordsWaterfallMedicineClinical endpointRandomized controlled trialSample size determinationClinical trialStatisticsArtificial intelligenceInternal medicineComputer scienceMathematics

Abstract

fetched live from OpenAlex

107 Background: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is an emerging short-term endpoint that may represent a surrogate for survival-based outcomes such as PFS and OS. We hypothesize that the configuration of waterfall plots in randomized trials may predict PFS/OS outcomes. Methods: A literature-based search was performed for all phase II/III randomized clinical trials published in MEDLINE from 2010 to 2022 testing molecularly targeted agents (MTA) or immunotherapy (IO). Articles reporting at least 1 waterfall plot for each treatment arm depicting maximum DepOR of target lesions with corresponding PFS/OS Kaplan-Meier plots were included. Studies are defined as positive or negative based on the achievement of a priori stated primary endpoint. Trial data collected included sample size per arm, cancer type, mechanisms of action of drug(s) tested, line of treatment, etc. Images of waterfall plots were manually extracted from publications and then processed through a semi-automatic extraction process using WebPlotDigitizer and Tesseract to produce tabular representations. Logistic regression with L2 regularization was used for modeling; hyperparameter tuning was accomplished with five-fold cross-validation on a training set compromising 80% of the data. Results: A total of 111 studies were identified: 65 (59%) phase III and 46 (41%) phase II, mean sample size per arm 317 (19-1581). Most frequent cancer type was gastrointestinal 24 (22%). MTA, IO and combinations were tested in 113 (51%), 35(16%) and 11 (5%) studies respectively. Chemotherapy and other treatment regimens were used in 63 (28%) trials. PFS was the primary endpoint in 62 (56%); 80 (75%) studies were positive. Of the 111 studies only 83 (75%) were retained for machine learning analysis, the remainder were excluded due to atypical formatting such as superimposed waterfall plots. Performance of the model was assessed on a test set which comprised 20% of the original dataset. Table below shows the classification metrics from modelling. Both PPV and NPV were ≥80%. Conclusions: MAP-OUTCOMES evaluated pan-cancer randomized studies with diverse therapeutic anticancer agents. It is a computational tool with the potential to predict survival-based outcomes from waterfall plots and may help with decisions regarding follow-on randomized studies. Further validation is ongoing. [Table: see text]

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.091
metaresearch head score (Gemma)0.238
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.091
Threshold uncertainty score0.480

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0910.238
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0030.006
Bibliometrics0.0100.005
Science and technology studies0.0000.001
Scholarly communication0.0040.004
Open science0.0020.002
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0060.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.180
GPT teacher head0.445
Teacher spread0.265 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJCO Global OncologySame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207