The use of machine learning models to predict progression-free survival and overall survival outcomes from waterfall plots in randomized clinical trials (MAP-OUTCOMES)
Notice bibliographique
Résumé
BACKGROUND: Depth of tumor response (DepOR) of individual patients, as visualized by waterfall plots, is a short-term endpoint that may represent a surrogate for survival-based outcomes such as progression-free survival (PFS) and overall survival (OS). We hypothesized that PFS/OS could be predicted from waterfall plots in randomized clinical trials (RCTs) using a novel machine-learning (ML) computational model. MATERIALS AND METHODS: A literature-based search was carried out for phase II/III RCTs testing noncytotoxic systemic therapy, which included waterfall plots with corresponding PFS/OS results. Studies were defined as positive or negative based on achievement of an a priori-stated primary endpoint. Trial data and images of waterfall plots were manually extracted and then processed through a semi-automatic extraction process. We developed the MAP-OUTCOMES (MAchine learning model to Predict PFS and OS OUTCOMES) model using regularized logistic regression. This model was applied to a training set comprising 70% of the data, and 30% was used for a test set. RESULTS: A total of 91 unique RCTs were identified, and 82 (93 trial pairs) retained for the ML analysis. Most of the trials were phase III (75%), with 67% using PFS as the primary endpoint and a mean sample size of 350 patients per arm. The most common tumor type was genitourinary (22%), and small-molecule targeted agents (27%) were the most frequent regimen. The model's performance achieved 71% accuracy [95% confidence interval (CI) 0.536-0.862, P = 0.18] with an area under the curve (AUC) of 65% (95% CI 0.333-0.938, P = 0.157) and area under the precision-recall curve (AUPRC) of 90% (95% CI 0.779-0.995, P = 0.171) in the 28 trials used for the test set. CONCLUSIONS: The MAP-OUTCOMES model demonstrated the feasibility of using ML to predict survival-based outcomes from waterfall plots, thus providing a potential tool for early trial evaluation. Improving the model's performance with more training data and creating independent datasets are necessary steps to assess its generalizability for prospective clinical applications.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,025 | 0,046 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,003 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,000 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».