MétaCan
Menu
Back to cohort
Record W4408166624 · doi:10.2196/66821

Augmenting Insufficiently Accruing Oncology Clinical Trials Using Generative Models: Validation Study

2025· article· en· W4408166624 on OpenAlexaff
Samer El Kababji, Nicholas Mitsakakis, Elizabeth Jonker, Ana-Alicia Beltran-Bless, Gregory R. Pond, Lisa Vandermeer, Dhenuka Radhakrishnan, Lucy Mosquera, Alexander Paterson, Lois E. Shepherd, Bingshu E. Chen, William E. Barlow, Julie R. Gralow, Marie-France Savard, Christian Fesl, Dominik Hlauschek, Marija Balić, Gabriel Rinnerthaler, Richard Greil, Michael Gnant, Mark Clemons, Khaled El Emam

Bibliographic record

VenueJournal of Medical Internet Research · 2025
Typearticle
Languageen
FieldMedicine
TopicEthics in Clinical Research
Canadian institutionsQueen's UniversityAlberta Health ServicesMcMaster UniversityAgricultural Research Institute of OntarioUniversity of Ottawa
FundersNational Cancer Institute
KeywordsMedicineClinical trialMedical physicsComputer scienceInternal medicine

Abstract

fetched live from OpenAlex

BACKGROUND: Insufficient patient accrual is a major challenge in clinical trials and can result in underpowered studies, as well as exposing study participants to toxicity and additional costs, with limited scientific benefit. Real-world data can provide external controls, but insufficient accrual affects all arms of a study, not just controls. Studies that used generative models to simulate more patients were limited in the accrual scenarios considered, replicability criteria, number of generative models, and number of clinical trials evaluated. OBJECTIVE: This study aimed to perform a comprehensive evaluation on the extent generative models can be used to simulate additional patients to compensate for insufficient accrual in clinical trials. METHODS: We performed a retrospective analysis using 10 datasets from 9 fully accrued, completed, and published cancer trials. For each trial, we removed the latest recruited patients (from 10% to 50%), trained a generative model on the remaining patients, and simulated additional patients to replace the removed ones using the generative model to augment the available data. We then replicated the published analysis on this augmented dataset to determine if the findings remained the same. Four different generative models were evaluated: sequential synthesis with decision trees, Bayesian network, generative adversarial network, and a variational autoencoder. These generative models were compared to sampling with replacement (ie, bootstrap) as a simple alternative. Replication of the published analyses used 4 metrics: decision agreement, estimate agreement, standardized difference, and CI overlap. RESULTS: Sequential synthesis performed well on the 4 replication metrics for the removal of up to 40% of the last recruited patients (decision agreement: 88% to 100% across datasets, estimate agreement: 100%, cannot reject standardized difference null hypothesis: 100%, and CI overlap: 0.8-0.92). Sampling with replacement was the next most effective approach, with decision agreement varying from 78% to 89% across all datasets. There was no evidence of a monotonic relationship in the estimated effect size with recruitment order across these studies. This suggests that patients recruited earlier in a trial were not systematically different than those recruited later, at least partially explaining why generative models trained on early data can effectively simulate patients recruited later in a trial. The fidelity of the generated data relative to the training data on the Hellinger distance was high in all cases. CONCLUSIONS: For an oncology study with insufficient accrual with as few as 60% of target recruitment, sequential synthesis can enable the simulation of the full dataset had the study continued accruing patients and can be an alternative to drawing conclusions from an underpowered study. These results provide evidence demonstrating the potential for generative models to rescue poorly accruing clinical trials, but additional studies are needed to confirm these findings and to generalize them for other diseases.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.084
metaresearch head score (Gemma)0.197
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.916
Threshold uncertainty score0.447

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0840.197
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0010.004
Bibliometrics0.0010.001
Science and technology studies0.0000.002
Scholarly communication0.0010.001
Open science0.0020.002
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0020.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.928
GPT teacher head0.786
Teacher spread0.142 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designSimulation or modeling
DomainMethods
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Medical Internet ResearchSame topicEthics in Clinical ResearchFrench-language works237,207