Point/counterpoint: randomized versus single-arm phase II clinical trials for patients with newly diagnosed glioblastoma
Bibliographic record
Abstract
In this article, we attempt to delineate the pros and cons of single-arm versus randomized phase II studies in patients with newly diagnosed glioblastoma (GBM), taking into account such factors as (i) the availability of appropriate controls, (ii) the interpretability of the resulting data, (iii) the goal of rapidly screening many novel agents using as few patients as necessary, (iv) utilization of limited financial and patient resources, and (v) maximization of patient participation in these studies. Phase II trials are typically considered middle development studies and address questions related to clinical outcome and tolerability. The overall goal is to obtain preliminary estimates of the likelihood a patient will benefit from treatment as well as the likelihood a patient may suffer a serious adverse effect from treatment. This informs the potential risk-benefit evaluation of the drug upon which decisions are made to further evaluate the drug in a definitive phase III trial or to stop further study of the regimen. Traditionally, phase II drug studies in oncology assess adverse events using the Common Terminology Criteria for Adverse Events scale and usually do not have formal rules for the determination of whether the regimen is deemed too toxic; this is often left to clinical judgment of the acceptability of the adverse event profile and severity in the context of an estimate of the potential benefit; there is more tolerance of adverse events for regimens that potentially have greater clinical benefit. Early in the era of oncology drug development, there were few effective agents across all cancer types. The primary role of a phase II trial was to quickly screen out clinically ineffective agents while minimizing the number of patients exposed to them. Desirable properties of an endpoint in this setting are that it can be evaluated in a short time, it is suggestive of clinical benefit, and it is minimally impacted by patient and disease characteristics so that the mix of patients in the trial would have little influence on the trial results. The endpoint most commonly used that meets these criteria is tumor shrinkage. Early designs were meant to minimize the number of patients exposed to the tested drug, since it was likely to be ineffective and toxic. A drug with a tumor response rate less than 20% was felt to be not promising. The Gehan design1 was the first used. It is a 2-stage single-arm design in which an initial cohort of 14 patients is accrued. If no responses are observed, the drug is declared ineffective with 95% confidence the tumor response rate is less than 20%. If one or more responses are seen, an additional 11–16 patients are accrued to yield a response rate estimate for which the margin of error is at most 0.10 for the 2-sided 95% CI. The Gehan design provides limited guidance for determining whether an observed response rate is clinically meaningful and does not provide information regarding the probabilities of type I and type II errors. As more effective oncology drugs were becoming available, a higher standard of evidence for determining the potential for clinical benefit became important. Fleming2 proposed a 2-stage design that requires the specification of the smallest response rate that would be considered promising and the largest response rate that would be deemed not to be promising with specified type I and type II error probabilities. Stopping boundaries are developed for the first stage that recommend stopping if the results are drastically positive or negative. If the boundaries are not crossed at the first stage, additional patients are accrued and decision rules are applied to determine whether the drug has potential clinical benefit or not. Simon3 optimized this design with the modification that the trial would stop after the first stage only if the results are deemed dramatically negative. If the stage 1 results are promising, the additional stage 2 patients yield a more precise estimate of the response rate. This design minimizes the average sample size under the null hypothesis that the response rate is not promising and achieves the desired type I and type II error probabilities. The Simon design, or some variation, remains the standard for single-arm phase II trials.4–6 Recently the increased availability of combination and molecularly targeted therapies has challenged the use of a single-arm phase II trial design. The concern is the reliability of a historical control value. In combination therapies, presumably one or both treatments are effective to some degree and a precise estimate of tumor response rate for each monotherapy is not available. It is likely that molecularly targeted therapies exert disease control through mechanisms other than tumor shrinkage, and so tumor response is not an adequate endpoint. If progression-free survival (PFS) is used, historical estimates are unreliable because this endpoint is influenced by disease and patient characteristics, and the relatively small cohort of patients in the new trial may differ significantly from the historical controls (if available). Furthermore, there may not be a historical control at all if patients are selected on the basis of having the molecular target. The presence of the molecular target may be prognostic and when the molecular target status is not known in the historical control cohort, a reliable historical control value cannot be obtained. Unreliable historical controls have given rise to randomized phase II trial designs. These trials generally randomize patients to standard treatment (control arm) and one or more experimental arms. A formal comparison is made between the treatment arm(s) and the control arm. There are different variations of a randomized phase II trial, which include the seamless phase II/III trials,7,8 randomized discontinuation trials,9 and Bayesian outcome-adaptive randomization trials.10,11 Seamless phase II/III designs are being used more often because of the efficiencies gained in terms of protocol development and site contracting. Essentially, the protocol suspends accrual at the completion of the phase II portion and if the phase II trial is positive, it is reopened for phase III accrual. Patients in the phase II portion can be used as part of the phase III portion. Hence there is only the need to develop and get approval for a single protocol, and site contracting needs to be done only once. Another design of recent interest uses Bayesian outcome-adaptive randomization. In this design, the randomization scheme is re-weighted after each patient outcome is observed to more favor the arm with the better outcomes. Hence the next patient randomized has a higher chance of being randomized to the arm that has the best outcomes at that point. This trial design minimizes the number of patients who receive ineffective treatment (if one treatment is better than the others); however, it generally requires larger sample sizes than randomization schemes that have equal probabilities of randomizing among the arms. Finally, another design of interest is the phase II screening design.12 These often do not have a control arm but rather compare multiple experimental treatments to select the one that has the best outcomes to move forward into a phase III trial. Randomized phase II trials often use an intermediate endpoint that differs from the endpoint for a definitive phase III trial. For example, the phase II trial endpoint is PFS and the subsequent phase III endpoint is overall survival (OS). Randomized phase II trials sometimes also measure the PFS rate at a specified time point, such as 12 months. This aids in ensuring that the final analysis can occur soon after the last patient enrolled has been followed for the necessary time (eg, 12 months for PFS rate at 12 months) rather than having to wait until a specified number of events has been observed, but this has very modest impact on sample size. Randomized phase II trials generally have a large type I error (eg, one-sided 0.10) requiring a subsequent phase III trial with a more definitive type I error rate (eg, 2-sided 0.05). Allowing larger type I errors and increasing the type II error (reducing power) reduces the sample sizes needed for a randomized phase II trial. This tends to make randomized phase II trials feasible in terms of sample size requirements. In summary, the intent of the randomized phase II is to determine whether a treatment has potential for clinical benefit in a timely manner, that is, to inform a go/no-go decision, and not to provide definitive evidence. Overall, there are advantages and disadvantages associated with single-arm phase II trials and randomized phase II trials. The advantage of using one design rather than the other depends upon several factors. The remainder of this paper illustrates and discusses the advantages and disadvantages of these designs within the context of neuro-oncology trials. The specific value of randomization in phase II is linked to the endpoint of the trial in question, which in turn is dependent on the overall goals of the study. For this discussion, we will assume that the purpose of phase II is to elucidate some measure of biological activity, or “signal,” and to provide sufficient data to adequately inform good phase III go/no-go decision making. An alternative purpose of a phase II trial might be to focus solely on the “signal finding” aspects, but then we would still be left with designing another trial to make decisions about whether to move to phase III. All therapeutic development is associated with risk, but oncology drug development is more likely to fail in the later stages of development, and failures in phase III are common.13 Go/no-go decision making would be improved if phase II results could reliably estimate the probability of phase III success. Phase III trials in neuro-oncology typically use OS as an endpoint, so it is critical that the effects on the endpoints chosen in phase II have some ability to predict therapeutic effects on OS. For clinical trials evaluating novel therapies to treat GBM, several endpoints have demonstrated potential for false signals. Overall response rate has proven to have a poor association with OS14 and effects on PFS correlate strongly with OS effects for temozolomide,14,15 but this association does not hold for bevacizumab and might be expected to vary with immunotherapy.16 Additionally, for a disease with no proven efficacious therapies and short survival time in the post-progression phase, it is questionable whether PFS provides meaningful benefit over OS as a trial endpoint in GBM.17 Regardless of the endpoint chosen, randomization has utility in separating therapeutic signal from confounders, and in the case of both PFS and OS, randomization is crucial. The first published randomized controlled trial was designed by Sir Austin Bradford Hill for the Medical Research Council to determine the value of streptomycin in treating pulmonary tuberculosis.18 Randomization had been strongly advocated in controlled experimentation by R. A. Fisher as a means to accurately estimate sampling error and legitimize significance testing.19 Hill additionally argued that randomization was important in clinical trials to control for the potential for selection bias resulting in observed outcomes that were attributable to factors other than the experimental intervention.19 Endpoints with significant natural variability and many potential explanatory variables associated with that variability are more prone to such bias. For example, OS varies substantially among patients and may be attributable to known prognostic factors such as age and performance status in addition to potential unmeasured confounders. Alternatively, overall response rate may be more directly attributable to therapy with fewer alternative explanatory variables. For any endpoint, randomization is used to control for such confounders and improve phase II trial design,20 but it is particularly important for endpoints such as PFS and OS. In addition to selection bias, comparison to historical controls may be prone to false positive results by ignoring the variability in the historical control and not accounting for patient temporal drift.21 Comparison of single-arm phase II results to historical data can therefore lead to an overestimation of therapeutic effect and result in poor phase III go/no-go decision making and late stage failures. Maitland et al22 showed that the overall predictive probability of phase II “success” in combination chemotherapy trials is extremely low and that phase II trials were more likely to claim success if they were not randomized. Simulation studies have demonstrated that while single-arm and randomized studies unsurprisingly are comparable as long as there is a strong historical control for a given endpoint,23 the addition of selection bias or patient temporal drift can result in significantly higher false positive rates than randomized studies.23 Patient selection and temporal drift in real world single-arm studies are generally concerns in clinical trials and have been demonstrated in neuro-oncology, specifically. Grossman et al21 published the results of 3 separate single-arm trials with diverse mechanisms of action conducted through the New Approaches to Brain Tumor Therapy (NABTT) in comparison with historical data from both NABTT and the European Organisation for Research and Treatment of Cancer (EORTC)/National Cancer Institute of Canada (NCIC) study that defined standard of care. Even though known prognostic factors were similar or of slightly higher risk in the single-arm trials, all 3 trials showed substantially better OS compared with the EORTC/NCIC study but notably also compared with the more internally standardized historical data from NABTT.21 The likelihood of having 3 effective drugs in a disease with so few historical successes is low and is more easily explained by the aforementioned problems with historical controls, even when attempts are made to control for the known issues. One of the drugs included in this analysis, cilengitide, had another single-arm phase II trial that was interpreted as promising compared with historical controls.24 Unfortunately, the optimism was not confirmed in CENTRIC, the follow-up randomized phase III study.25 Similarly, the ACT IV trial26 of rindopepimut in addition to standard chemoradiotherapy for newly diagnosed GBM failed to validate the excitement generated from 3 uncontrolled phase II studies.27–29 The most common criticism for incorporating randomization into phase II trials is based on efficiency. To add randomization requires the addition of a control arm and the variability in outcome associated with that arm, leading to a substantial increase in total number of patients required for the trial. This argument is even more compelling in GBM, where randomization to a standard of care with such poor outcomes is undesirable. Reliable data in the phase II setting prevent additional patients from becoming exposed to ineffective therapies in phase III, however, and there are mechanisms to mitigate the increased patient resources required for randomization. Unequal randomization strategies, such as 2:1, have been employed to be more attractive to patients, but anything without equal numbers of control patients to a given experimental arm results in overall efficiency loss in the trial. Randomized, noncomparative studies ultimately still rely on historical control estimates. Another solution is to have control arms that are common to multiple experimental therapies in platform trials under master protocols. When multiple arms are used in a single clinical trial infrastructure, response adaptive randomization can be added to more efficiently allocate patients to successful arms. Even R. A. Fisher is reported to have suggested that randomization proportions may need to be dynamically altered based on accumulating results in medical trials.19 Such a design is most notably displayed in practice in the ongoing I-SPY 2 trial for neoadjuvant systemic therapy in breast cancer.30 Translating such elements to GBM31 could provide efficiencies for phase II trials based on OS while maintaining 1:1 comparison between promising arms and control.32 The National Cancer Institute’s Brain Malignancy Steering Committee Clinical Trials Planning Workshop included these elements in published recommendations33 and they are now being included in the recently opened INdividualized Screening trial of Innovative Glioblastoma Therapy (INSIGhT; NCT02977780) and the GBM Adaptive Global Learning Environment (GBM AGILE) trial, currently in development. In reality, the standard paradigm of phase I → phase II → phase III is too simplistic to describe actual trial goals. For early phase II trials where a change in an imaging or pharmacodynamic biomarker is anticipated based on mechanism of action, a single-arm study may be appropriate to investigate biological and determine whether further study is These may also be from phase I however, so the between are decisions need to be made as to whether an experimental therapy has potential to improve OS to the of large patient and financial resources in studies. For all of the randomization is critical in reliable data to make particularly for clinical trials in diagnosed GBM is a survival and the average is months in patients who are well to clinical and less than 12 months in the of and the accrual of over patients to only 2 drugs and have been by the for patients with newly diagnosed Unfortunately, these survival by less than 3 and a potential for The to substantial include (i) an to this (ii) the of the to (iii) of the of chemotherapy to the and (iv) the of to This of substantial in the treatment of newly diagnosed the need to screen a of novel for preliminary evidence of using phase II trials. This is more than one might in this patient the number of patients and to these trials is GBM is a relatively with new each in the of patients with this disease are over of only of all cancer patients in clinical trials, and the at a substantially factors that can participation in clinical trials for newly diagnosed GBM include a of that may a ability to to or treatments and that may with the ability to provide A to it is to assess the of trial in patients with newly diagnosed response by is with error as may not even as tumor Similarly, may improve with therapies that treat or with that of the status of the endpoints are also often as in and provides information on the of rather than the size of the As a these typically better after treatment with or therapies and of the status of the The between and outcome has recently been in phase III trials of bevacizumab which on but had no impact on These the use of response such as OS as the best endpoint for phase II studies in this The goal of phase II trials in newly diagnosed is to rapidly screen new agents and for an early signal using as few patients as In an world with few on patients and resources, randomized phase II trials would be the as they significantly the false positive rate by bias and increasing confidence in the of drug a recent of phase II oncology clinical trials that to phase III studies failed to a significant between single-arm and randomized phase II designs in the phase III given the long of trials in patients with GBM, the primary the design of phase II studies might favor designs on ineffective therapies for early phase rather than designs that each new drug might be efficacious phase Furthermore, as patients for these trials are other of phase II trial designs Randomized phase II trials significantly higher patient numbers study to the of an control arm which standard This design also the time to study completion and the the number of novel agents that can be Finally, many and patients who are and to in new drug studies have a strong for phase II studies also controls with similar criteria are required for the results of these studies to be the newly diagnosed GBM currently randomized phase III studies with control arms which can provide outcome when a historical control is not available, a randomized phase II design be For example, when molecularly defined are selected for the natural of these are usually and a control arm would be necessary to accurately determine the impact of any In the most and reliable endpoints be used when a single-arm study to minimize variables. In newly diagnosed GBM trials the most reliable endpoint is OS. An appropriate concern in using historical controls for GBM trials is that patient survival can improve over time as more to do multiple patients with and better For this historical controls are the outcomes of clinical studies in with GBM over the several it is relatively to that the of designed single-arm phase II studies evaluating novel therapies will fail to a survival signal and would not be for further development. An result from a single-arm study be considered only as a preliminary signal of and be followed by a randomized phase II study with or without to to a phase III Phase II trials are important in the development of novel therapies for patients with newly diagnosed These trials the initial evaluation of in this disease and are designed to novel and into that do and do not further study. it often that there is a between randomized and single-arm phase II trial in each has in effective drug development. the of experimental drugs tested in patients with newly diagnosed GBM have been clinically designing small single-arm phase II studies to ineffective therapies early is larger randomized phase II trials are important to confounders and false the trial endpoint, the availability of appropriate controls, and trial results will be used to inform further development of the experimental therapy when a design. single-arm and randomized phase II trials well as adaptive and provide for evaluating the of novel in newly diagnosed glioblastoma when applied in an appropriate This not rely on any of interest for any
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.027 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".