The use of machine learning models to predict PFS and OS outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES).
Bibliographic record
Abstract
107 Background: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is an emerging short-term endpoint that may represent a surrogate for survival-based outcomes such as PFS and OS. We hypothesize that the configuration of waterfall plots in randomized trials may predict PFS/OS outcomes. Methods: A literature-based search was performed for all phase II/III randomized clinical trials published in MEDLINE from 2010 to 2022 testing molecularly targeted agents (MTA) or immunotherapy (IO). Articles reporting at least 1 waterfall plot for each treatment arm depicting maximum DepOR of target lesions with corresponding PFS/OS Kaplan-Meier plots were included. Studies are defined as positive or negative based on the achievement of a priori stated primary endpoint. Trial data collected included sample size per arm, cancer type, mechanisms of action of drug(s) tested, line of treatment, etc. Images of waterfall plots were manually extracted from publications and then processed through a semi-automatic extraction process using WebPlotDigitizer and Tesseract to produce tabular representations. Logistic regression with L2 regularization was used for modeling; hyperparameter tuning was accomplished with five-fold cross-validation on a training set compromising 80% of the data. Results: A total of 111 studies were identified: 65 (59%) phase III and 46 (41%) phase II, mean sample size per arm 317 (19-1581). Most frequent cancer type was gastrointestinal 24 (22%). MTA, IO and combinations were tested in 113 (51%), 35(16%) and 11 (5%) studies respectively. Chemotherapy and other treatment regimens were used in 63 (28%) trials. PFS was the primary endpoint in 62 (56%); 80 (75%) studies were positive. Of the 111 studies only 83 (75%) were retained for machine learning analysis, the remainder were excluded due to atypical formatting such as superimposed waterfall plots. Performance of the model was assessed on a test set which comprised 20% of the original dataset. Table below shows the classification metrics from modelling. Both PPV and NPV were ≥80%. Conclusions: MAP-OUTCOMES evaluated pan-cancer randomized studies with diverse therapeutic anticancer agents. It is a computational tool with the potential to predict survival-based outcomes from waterfall plots and may help with decisions regarding follow-on randomized studies. Further validation is ongoing. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.035 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".