Systematic review of basket trials, umbrella trials, and platform trials: a landscape analysis of master protocols
Bibliographic record
Abstract
BACKGROUND: Master protocols, classified as basket trials, umbrella trials, and platform trials, are novel designs that investigate multiple hypotheses through concurrent sub-studies (e.g., multiple treatments or populations or that allow adding/removing arms during the trial), offering enhanced efficiency and a more ethical approach to trial evaluation. Despite the many advantages of these designs, they are infrequently used. METHODS: We conducted a landscape analysis of master protocols using a systematic literature search to determine what trials have been conducted and proposed for an overall goal of improving the literacy in this emerging concept. On July 8, 2019, English-language studies were identified from MEDLINE, EMBASE, and CENTRAL databases and hand searches of published reviews and registries. RESULTS: We identified 83 master protocols (49 basket, 18 umbrella, and 16 platform trials). The number of master protocols has increased rapidly over the last five years. Most have been conducted in the US (n = 44/83) and investigated experimental drugs (n = 82/83) in the field of oncology (n = 76/83). The majority of basket trials were exploratory (i.e., phase I/II; n = 47/49) and not randomized (n = 44/49), and more than half (n = 28/48) investigated only a single intervention. The median sample size of basket trials was 205 participants (interquartile range, Q3-Q1 [IQR]: 500-90 = 410), and the median study duration was 22.3 (IQR: 74.1-42.9 = 31.1) months. Similar to basket trials, most umbrella trials were exploratory (n = 16/18), but the use of randomization was more common (n = 8/18). The median sample size of umbrella trials was 346 participants (IQR: 565-252 = 313), and the median study duration was 60.9 (IQR: 81.3-46.9 = 34.4) months. The median number of interventions investigated in umbrella trials was 5 (IQR: 6-4 = 2). The majority of platform trials were randomized (n = 15/16), and phase III investigation (n = 7/15; one did not report information on phase) was more common in platform trials with four of them using seamless II/III design. The median sample size was 892 (IQR: 1835-255 = 1580), and the median study duration was 58.9 (IQR: 101.3-36.9 = 64.4) months. CONCLUSIONS: We anticipate that the number of master protocols will continue to increase at a rapid pace over the upcoming decades. More efforts to improve awareness and training are needed to apply these innovative trial design methods to fields outside of oncology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.221 | 0.519 |
| Meta-epidemiology (narrow) | 0.003 | 0.002 |
| Meta-epidemiology (broad) | 0.015 | 0.016 |
| Bibliometrics | 0.034 | 0.040 |
| Science and technology studies | 0.002 | 0.003 |
| Scholarly communication | 0.008 | 0.011 |
| Open science | 0.004 | 0.006 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".