MétaCan
Menu
Retour à la cohorte
Enregistrement W4412440980 · doi:10.1016/j.esmoop.2025.105509

The use of machine learning models to predict progression-free survival and overall survival outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES)

2025· article· en· W4412440980 sur OpenAlexaff
Khadjah Alshankati, Aisha Alshibany, Azhar Toma, Katherine Lajkosz, Benjamin Haibe‐Kains, Lillian L. Siu

Notice bibliographique

RevueESMO Open · 2025
Typearticle
Langueen
DomaineMedicine
ThématiqueRadiomics and Machine Learning in Medical Imaging
Établissements canadiensVector InstitutePrincess Margaret Cancer CentreUniversity Health Network
Organismes subventionnairesnon disponible
Mots-clésWaterfallClinical endpointConfidence intervalMedicineRandomized controlled trialProgression-free survivalLogistic regressionWaterfall modelArtificial intelligenceMachine learningInternal medicineStatisticsOncologyOverall survivalComputer scienceMathematicsSoftwareCartography

Résumé

récupéré en direct d'OpenAlex

BACKGROUND: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is a short-term endpoint that may represent a surrogate for survival-based outcomes such as progression-free survival (PFS) and overall survival (OS). We hypothesized that PFS/OS could be predicted from waterfall plots in randomized clinical trials (RCTs) using a novel machine-learning (ML) computational model. MATERIALS AND METHODS: A literature-based search was carried out for phase II/III RCTs testing noncytotoxic systemic therapy, which included waterfall plots with corresponding PFS/OS results. Studies were defined as positive or negative based on achievement of an a priori-stated primary endpoint. Trial data and images of waterfall plots were manually extracted and then processed through a semi-automatic extraction process. We developed the MAP-OUTCOMES (MAchine learning model to Predict PFS and OS OUTCOMES) model using regularized logistic regression. This model was applied to a training set comprising 70% of the data, and 30% was used for a test set. RESULTS: A total of 91 unique RCTs were identified, and 82 (93 trial pairs) retained for the ML analysis. Most of the trials were phase III (75%), with 67% using PFS as the primary endpoint and a mean sample size of 350 patients per arm. The most common tumor type was genitourinary (22%), and small-molecule targeted agents (27%) were the most frequent regimen. The model's performance achieved 71% accuracy [95% confidence interval (CI) 0.536-0.862, P = 0.18] with an area under the curve (AUC) of 65% (95% CI 0.333-0.938, P = 0.157) and area under the precision-recall curve (AUPRC) of 90% (95% CI 0.779-0.995, P = 0.171) in the 28 trials used for the test set. CONCLUSIONS: The MAP-OUTCOMES model demonstrated the feasibility of using ML to predict survival-based outcomes from waterfall plots, thus providing a potential tool for early trial evaluation. Improving the model's performance with more training data and creating independent datasets are necessary steps to assess its generalizability for prospective clinical applications.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction distillée sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.

score de la tête « metaresearch » (Codex)0,025
score de la tête « metaresearch » (Gemma)0,046
Version: codex-gemma-dda1882f352aStatut de validation: machine_predicted_unvalidated
Catégories candidatesMétarecherche
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Observationnel · Signal consensuel: Observationnel
GenreSignal candidat: Empirique · Signal consensuel: Empirique
Score de désaccord entre enseignants0,514
Score d'incertitude au seuil0,996

Scores Codex et Gemma par catégorie

CatégorieCodexGemma
Métarecherche0,0250,046
Méta-épidémiologie (sens strict)0,0000,000
Méta-épidémiologie (sens large)0,0030,000
Bibliométrie0,0000,000
Études des sciences et des technologies0,0000,000
Communication savante0,0000,000
Science ouverte0,0010,001
Intégrité de la recherche0,0000,001
Charge utile insuffisante (le modèle a refusé de juger)0,0000,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,147
Tête enseignante GPT0,428
Écart entre enseignants0,281 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.

Devis d'étudeObservationnel
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations0
Publié2025
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueESMO OpenMême sujetRadiomics and Machine Learning in Medical ImagingTravaux en français237 207