Response: Re: Dose Escalation Methods in Phase I Cancer Clinical Trials
Bibliographic record
Abstract
We thank Zohar and O'Quigley for their interest in our review on dose escalation methods in phase I cancer clinical trials ( 1 ). They communicated concern about the lack of safety using the standard design. They also declared that the continual reassessment method (CRM) outperforms the standard method by more accurately defining the true maximum tolerated dose and by treating more patients at optimal dose levels. The first point pertains to the limitations of the traditional 3 + 3 method. We agree that existing dose escalation methods, which include but are not limited to the traditional 3 + 3 method, have their respective strengths and weaknesses. In fact, we highlighted the advantages and drawbacks of various dose escalations methods in table 2 of our review ( 1 ) so that these methods can be better appreciated and thus applied more appropriately in future phase I clinical trials. Zohar and O'Quigley claimed that the traditional 3 + 3 design makes inaccurate dose recommendations, but their argument is based primarily on statistical simulations or reanalysis of data from completed trials. To demonstrate that the traditional 3 + 3 design has not been unsafe in clinical practice, we reviewed anticancer agents that have received US Food and Drug Administration (FDA) approval [see table 3 in ( 1 )]. To ensure the comparability of dosing regimens, we retrieved published phase I trials of these agents that have used the same or similar dosing schedules as the FDA-approved schedules and compared the FDA-approved dose (or dose intensity if the schedules were slightly different) of each agent with the recommended phase II dose from the corresponding phase I trial ( Table 1 ). Among 20 phase I trials that used the traditional 3 + 3 design with or without intrapatient dose escalation, five (25%) established a recommended phase II dose that was higher than the FDA-approved dose by an average of 26% (range = 15%–43%). Of two phase I trials that used a modified CRM design, one did not determine the maximum tolerated dose and thus recommended further dose testing in phase II trials, whereas the other predicted a 20% higher recommended phase II dose than the FDA-approved dose. These data indicate that the traditional 3 + 3 design and the CRM-based designs do not differ from each other in terms of safety and accuracy as much as was suggested by Zohar and O'Quigley. In addition, the traditional 3 + 3 design is useful in certain settings, such as the development of cytotoxic agents with a targeted toxicity level between 20% and 30% ( 2 ). Food and Drug Administration (FDA)–approved doses and recommended phase II doses established in the phase I trials using the same or similar dosing schedules for recent anticancer agents * → = followed by; ATD = accelerated titration design; bid = twice a day; CRM = continual reassessment method; d × 14 = daily for 14 days; d × 4w = daily for 4 weeks; IPDE = intrapatient dose escalation; Mab = monoclonal antibody; NA = not applicable; PK = pharmacokinetic data; q2w = every 2 weeks; q3w = every 3 weeks; q6w = every 6 weeks; qw = weekly; qw × 3 q4w = weekly for 3 weeks every 4 weeks; qw × 4 q6w = weekly for 4 weeks every 6 weeks; qw × 7 = weekly for 7 weeks; STKI = serine and/or threonine kinase inhibitor; TKI = tyrosine kinase inhibitor. Of two trials that used modified CRM designs, one did not determine the maximum tolerated dose and the other had a recommended phase II dose that exceeded the FDA-approved dose. Phase I trials using traditional 3 + 3 designs with or without IPDE in which the recommended phase II dose exceeded the FDA-approved dose. The second point relates to the perceived superiority of CRM-based designs, which is based on simulation studies. We agree that the CRM-based methods have many theoretical advantages over the traditional 3 + 3 design; however, reality has shown that although they exposed fewer patients at earlier, potentially subtherapeutic doses, CRM-based designs have not resulted in shorter trial durations ( 3 , 4 ). To our knowledge, no empirical evidence exists to support the claim that CRM-based designs fare better than the traditional design in complex situations, including intrapatient dose escalation, drug combinations, or the inclusion of different subgroups in a single trial. We agree that new dose escalation methods are underused ( 5 ) and have encouraged the use and evaluation of innovative designs ( 1 ). New dose-ranging methodology is especially needed for molecularly targeted agents, which may have toxicity and efficacy profiles that differ from those of conventional cytotoxic agents ( 6 ). Although statistical simulations are informative in the assessment of new dose escalation methods, their clinical utility can only be validated through practical applications. After all, “the proof of the pudding is in the eating.”
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.097 | 0.504 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".