Author response: A meta-analysis of threats to valid clinical inference in preclinical research of sunitinib
Notice bibliographique
Résumé
Article Figures and data Abstract eLife digest Introduction Results Discussion Materials and methods References Decision letter Author response Article and author information Metrics Abstract Poor study methodology leads to biased measurement of treatment effects in preclinical research. We used available sunitinib preclinical studies to evaluate relationships between study design and experimental tumor volume effect sizes. We identified published animal efficacy experiments where sunitinib monotherapy was tested for effects on tumor volume. Effect sizes were extracted alongside experimental design elements addressing threats to valid clinical inference. Reported use of practices to address internal validity threats was limited, with no experiments using blinded outcome assessment. Most malignancies were tested in one model only, raising concerns about external validity. We calculate a 45% overestimate of effect size across all malignancies due to potential publication bias. Pooled effect sizes for specific malignancies did not show apparent relationships with effect sizes in clinical trials, and we were unable to detect dose–response relationships. Design and reporting standards represent an opportunity for improving clinical inference. https://doi.org/10.7554/eLife.08351.001 eLife digest Developing a new drug can take years, partly because preclinical research on non-human animals is required before any clinical trials with humans can take place. Nevertheless, only a fraction of cancer drugs that are put into clinical trials after showing promising results in preclinical animal studies end up proving safe and effective in human beings. Many researchers and commentators have suggested that this high failure rate reflects flaws in the way preclinical studies in cancer are designed and reported. Now, Henderson et al. have looked at all the published animal studies of a cancer drug called sunitinib and asked how well the design of these studies attempted to limit bias and match the clinical scenarios they were intended to represent. This systematic review and meta-analysis revealed that many common practices, like randomization, were rarely implemented. None of the published studies used 'blinding', whereby information about which animals are receiving the drug and which animals are receiving the control is kept from the experimenter, until after the test; this technique can help prevent any expectations or personal preferences from biasing the results. Furthermore, most tumors were tested in only one model system, namely, mice that had been injected with specific human cancer cells. This makes it difficult to rule out that any anti-cancer activity was in fact unique to that single model. Henderson et al. went on to find evidence that suggests that the anti-cancer effects of sunitinib might have been overestimated by as much as 45% because those studies that found no or little anti-cancer effect were simply not published. Though it is known that the anti-cancer activity of the drug increases with the dose given in both human beings and animals, an evaluation of the effects of all the published studies combined did not detect such a dose-dependent response. The poor design and reporting issues identified provide further grounds for concern about the value of many preclinical experiments in cancer. These findings also suggest that there are many opportunities for improving the design and reliability of study reports. Researchers studying certain medical conditions (such as strokes) have already developed, and now routinely implement, a set of standards for the design and reporting of preclinical research. It now appears that the cancer research community should do the same. https://doi.org/10.7554/eLife.08351.002 Introduction Preclinical experiments provide evidence of clinical promise, inform trial design, and establish the ethical basis for exposing patients to a new substance. However, preclinical research is plagued by poor design and reporting practices (van der Worp et al., 2010; Begley, 2013a; Begley and Ioannidis, 2015). Recent reports also suggest that many effects in preclinical studies fail replication (Begley and Ellis, 2012). Drug development efforts grounded on non-reproducible findings expose patients to harmful and inactive agents; they also absorb scarce scientific and human resources, the costs of which are reflected as higher drug prices. Several studies have evaluated the predictive value of animal models in cancer drug development (Johnson et al., 2001; Voskoglou-Nomikos et al., 2003; Corpet and Pierre, 2005). However, few have systematically examined experimental design—as opposed to use of specific models—and its impact on effect sizes across different malignancies (Amarasingh et al., 2009; Hirst et al., 2013). A recent systematic review of guidelines for limiting bias in preclinical research design was unable to identify any guidelines in oncology (Henderson et al., 2013). Validity threats in preclinical oncology may be particularly important to address in light of the fact that cancer drug development has one of the highest rates of attrition (Hay et al., 2014), and oncology drug development commands billions of dollars in funding each year (Adams and Brantner, 2006). In what follows, we conducted a systematic review and meta-analysis of features of design and outcomes for preclinical efficacy studies of the highly successful drug sunitinib. Sunitinib is a multi-targeted tyrosine kinase inhibitor sunitinib (SU11248, Sutent) and is licensed as monotherapy for three different malignancies (Chow and Eckhardt, 2007; Raymond et al., 2011). As it was introduced into clinical development around 2000 and tested against numerous malignancies, sunitinib provided an opportunity to study a large sample of preclinical studies across a broad range of malignancies—including several supporting successful translation trajectories. Results Study characteristics Our screen from database and reference searches captured 74 studies eligible for extraction, corresponding to 332 unique experiments investigating tumor volume response (Figure 1, Table 1, Table 1—source data 1E). Effect sizes (standardized mean difference [SMD] using Hedges' g) could not be computed for 174 experiments (52%) due to inadequate reporting (e.g., sample size not provided, effect size reported as a median, lack of error bars, Figure 1—figure supplement 1). Overall, 158 experiments, involving 2716 animals, were eligible for meta-analysis. The overall pooled SMD for all extracted experiments across all malignancies was −1.8 [−2.1, −1.6] (Figure 2—figure supplement 1). Mean duration of experiments used in meta-analysis (Figures 2–4) was 31 days (±14 days standardized deviation of the mean (SDM)). Figure 1 with 1 supplement see all Download asset Open asset Descriptive analysis of (A) internal, construct, and (B) external validity design elements. External validity scores were calculated for each malignancy type tested, according to the formula: number species used + number of models used; an extra point was assigned if a malignancy type tested more than one species and more than one model. https://doi.org/10.7554/eLife.08351.003 Figure 1—source data 1 (A) Coding details for IV and CV categories. https://doi.org/10.7554/eLife.08351.004 Download elife-08351-fig1-data1-v1.docx Table 1 Demographics of included studies https://doi.org/10.7554/eLife.08351.006 Study level demographicsIncluded studies (n = 74)Conflict of interest Declared19 (26%)Funding statement* Private, for-profit44 (59%) Private, not-for-profit35 (47%) Public37 (50%) Other2 (3%)Recommended clinical testing Yes37 (50%)Publication date 2003–200613 (18%) 2007–200917 (23%) 2010–201344 (59%) * Does not sum to 100% as many studies declared more than one funding source. Table 1—source data 1 (C) Search Strategies. (D) PRISMA Flow Diagram. (E) Demographics of included studies at qualitative level. https://doi.org/10.7554/eLife.08351.007 Download elife-08351-data1-v1.docx Figure 2 with 1 supplement see all Download asset Open asset Summary of pooled SMDs for each malignancy type. Shaded region denotes the pooled standardized mean difference (SMD) and 95% confidence interval (CI) (−1.8 [−2.1, −1.6]) for all experiments combined at the last common time point (LCT). https://doi.org/10.7554/eLife.08351.008 Figure 2—source data 1 (B) Heterogeneity statistics (I2) for each malignancy sub-group. https://doi.org/10.7554/eLife.08351.009 Download elife-08351-fig2-data1-v1.docx Figure 3 Download asset Open asset Relationship between study design elements and effect sizes. The shaded region denotes the pooled SMD and 95% CI (−1.8 [−2.1, −1.6]) for all experiments combined at the LCT. https://doi.org/10.7554/eLife.08351.011 Figure 4 Download asset Open asset Funnel plot to detect publication bias. Trim and fill analysis was performed on pooled malignancies, as well as the three malignancies with the greatest study volume. (A) All experiments for all malignancies (n = 182), (B) all experiments within renal cell carcinoma (RCC) (n = 35), (C) breast cancer (n = 32), and (D) colorectal cancer (n = 29). Time point was the LCT. Open circles denote original data points whereas black circles denote 'filled' experiments. Trim and fill did not produce an estimate in RCC; therefore, no overestimation of effect size could be found. https://doi.org/10.7554/eLife.08351.012 Design elements addressing validity threats Effects in preclinical studies can fail clinical generalization because of bias or random variation (internal validity), a mismatch between experimental operations and the clinical scenario modeled (construct validity), or idiosyncratic causal mediators in an experimental system (external validity) (Henderson et al., 2013). We extracted design elements addressing each using consensus design practices identified in a systematic review of validity threats in preclinical research (Henderson et al., 2013). Few studies used practices like blinding or randomization to address internal validity threats (Figure 1A). Only 6% of experiments investigated a dose–response relationship (3 or more doses). Concealment of allocation or blinded outcome assessment was never reported in studies that advanced to meta-analysis. It is worth noting that one research group employed concealed allocation and blinded assessment for the many experiments it described (Maris et al., 2008). However, statistics were reported in a way that did not align with those we needed to calculate SMD. We found that 58.8% of experiments included active drug comparators, thus, facilitating interpretation of sunitinib activity (however, we note that in some of the experiments, sunitinib was an active comparator in a test of a different drug or drug combination). Construct validity practices can only be meaningfully evaluated against a particular, matched clinical trial. Nevertheless, Figure 1A shows that experiments predominantly relied on juvenile, female, immunocompromised mouse models, and very few animal efficacy experiments used genetically engineered cancer models (n = 4) or spontaneously arising tumors (n = 0). Malignancies generally scored low (score = 1) for addressing external validity (Figure 1B), with breast cancer studies employing the greatest variety of species (n = 2) and models (n = 4). Implementation of internal validity practices did not show clear relationships with effect sizes (Figure 3A). However, sunitinib effect sizes were significantly greater when active drug comparators were present in an experiment compared to when they were not (−2.2 [−2.5, −1.9] vs −1.4 [−1.7, −1.1], p-value <0.001). Within construct validity, there was a significant difference in pooled effect size between genetically engineered mouse models and human xenograft (p-value <0.0001) and allograft (p-value 0.001) model types (Figure 3B). For external validity (Figure 3C), malignancies tested in more and diverse experimental systems tended to show less extreme effect sizes (p < 0.001). Evidence of publication bias For the 158 individual experiments, 65.8% showed statistically significant activity at the experiment level (p < 0.05, Figure 2—figure supplement 1), with an average sample size of 8.03 animals per treatment arm and 8.39 animals per control arm. Funnel plots for all studies (Figure 4A), as well as our renal cell carcinoma (RCC) subset (Figure 4B) suggest potential publication bias. Trim and fill analysis suggests an overestimation of effect size of 45% (SMD changed from −1.8 [−2.1, −1.7] to −1.3 [−1.5, −1.0]) across all indications. For high-grade glioma and breast cancer, the overestimation was 11% and 52%, respectively. However, trim and fill analysis suggested excellent symmetry for the RCC subgroup, suggesting coverage of the overall effect size and confidence intervals and not overestimation of effect size. Preclinical studies and clinical correlates Every malignancy tested with sunitinib showed statistically significant anti-tumor activity (Figure 2). Though we did not perform a systematic review to estimate clinical effect sizes for sunitinib against various malignancies, a perusal of the clinical literature suggests little relationship between pooled effect sizes and demonstrated clinical activity. For instance, sunitinib monotherapy is highly active in RCC patients (Motzer et al., 2006a, 2006b) and yet showed a relatively small preclinical effect; in contrast, sunitinib monotherapy was inactive against small cell lung cancer in a phase 2 trial (Han et al., 2013), but showed relatively large preclinical effects. Using measured effect sizes at a standardized time point of 14 days after first administration (a different time point than in Figures 2–4 to better align our evaluation of dose–response), we were unable to observe a dose–response relationship over three orders of magnitude (0.2–120 mg/kg/day) for all experiments (Figure 5A). We were also unable to detect a dose–response relationship over the full dose range (4–80 mg/kg/day) tested in the RCC subset (Figure 5B). The same results were observed when we performed the same analyses using the last time point in common between the experimental and control arms. Figure 5 Download asset Open asset Dose–response curves for sunitinib preclinical studies. Only experiments with a once daily (no breaks) administration schedule were included in both graphs. Effect size data were taken from a standardized time point (14 days after first sunitinib administration). (A) Experiments (n = 158) from all malignancies tested failed to show a dose–response relationship. (B) A dose–response relationship was not detected for RCC (n = 24). (C) Dose–response curves reported in individual studies within the RCC subset showed dose–response patterns (blue diamond = Huang 2010a [n = 3], red square = Huang 2010d [n = 3], green triangle = Ko 2010a [n = 3], purple X = Xin 2009 [n = 3]). https://doi.org/10.7554/eLife.08351.013 Discussion Preclinical studies serve an important role in formulating clinical hypotheses and justifying the advance of a new drug into clinical testing. Our meta-analysis, which included malignancies that respond to sunitinib in human beings and those that do not, raises several questions about methods and reporting practices in preclinical oncology—at least in the context of one well-established drug. First, reporting of design elements and data was poor and inconsistent with widely recognized standards for animal studies (Kilkenny et al., 2010). Indeed, 98 experiments (30% of qualitative sample) could not be quantitatively analyzed because sample sizes or measures of dispersion were not provided. Experimenters only sporadically addressed major internal validity threats and tended not to test indication-activity in more than one model and species. This finding is consistent with what others have observed in experimental stroke and other research areas (Macleod et al., 2004; van der Worp et al., 2005; Kilkenny et al., 2009; Glasziou et al., 2014). Some teams have shown a relationship between failure to address internal validity threats and exaggerated effect size (Crossley et al., 2008; Rooke et al., 2011); we did not observe a clear relationship. Consistent with what has been reported in stroke (O'Collins et al., 2006), our findings suggest that testing in more models tends to produce smaller effect sizes. However, since a larger sample of studies will provide a more precise estimate of effect, we cannot rule out that the trends observed for external validity reflect a regression to the mean. Second, preclinical studies for sunitinib seem to be prone to publication bias. Notwithstanding limitations on using funnel plots to detect publication bias (Lau et al., 2006), our plots were highly asymmetrical. That all malignancy types tested showed statistically significant anti-cancer activity strains credulity. Others have reported that far more animal studies report statistical significance than would be expected (Wallace et al., 2009; Tsilidis et al., 2013), and our observations that two thirds of individual studies showed significance extends these observations. Third, we were unable to detect a meaningful relationship between preclinical effect sizes and known clinical behavior. Although a full analysis correlating trial and preclinical effect sizes will be needed, we did not observe obvious relationships between the two. We also did not detect a dose–response effect over three orders of magnitude even within an indication—RCC—known to respond to sunitinib and even when different time points were used. It is possible that heterogeneity in cell lines or strains may have obscured the effects of dose. For example, experimenters may have delivered higher doses to xenografts known to show slow tumor growth. However, RCC patients—each of whom harbors genetically distinct tumors—show dose–response effects in trials (Faivre et al., 2006) and between trials in a meta-analysis (Houk et al., 2010). It is also possible that the toxicity of sunitinib may have limited the ability to demonstrate dose response, though this contradicts demonstration of dose response within studies (Abrams et al., 2003; Amino et al., 2006; Ko et al., 2010). Finally, the tendency for preclinical efficacy studies to report drug dose, but rarely drug exposure (i.e., serum measurement of active drug), further limits the construct validity of these studies (Peterson and Houghton, 2004). One explanation for our findings is that human xenograft models, which dominated our meta-analytic sample, have little predictive value, at least in the context of receptor tyrosine kinase inhibitors. This is a possibility that contradicts other reports (Kerbel, 2003; Voskoglou-Nomikos et al., 2003). We disfavor this explanation in light of the suggestion of publication bias; also, xenografts should show a dose–response regardless of whether they are useful clinical models. A second explanation is that experimental methods are so varied as to mask real effects. However, we note that the observed patterns on experimental design are based purely on what was reported in 'Materials and methods' section. Third, experiments assessing changes in tumor volume might only be interpretable in the context of other experiments within a preclinical report, such as with mechanistic and pharmacokinetic studies. This explanation is consistent with our observation that studies testing effect along a causal pathway tended to produce smaller effect sizes. A fourth possible explanation for our findings is that the predictive value of a small number of preclinical studies was obscured by inclusion of poorly designed and executed preclinical studies in our meta-analysis. Quantitative analysis of preclinical design factors that confer greater clinical generalizability awaits side-by-side comparison with pooled effects in clinical trials. Finally, it may be that design and reporting practices are so poor in preclinical cancer research as to make interpretation of tumor volume curves useless. Or, non-reporting may be so rampant as to render meta-analysis of preclinical research impossible. If so, this raises very troubling questions for the publication economy of cancer biology: even well-designed and reported studies may be difficult to interpret if their results cannot be compared to and synthesized with other studies. Our systematic review has several limitations. First, we relied on what authors reported in the published study. It is possible certain experimental practices, like randomization, were used but not reported in methods. Further to this, we relied only on published reports, and restriction of searches to the English language may have excluded some articles. In February of 2012, we filed a Freedom of Information Act from the and Drug for preclinical data in of 4 the has not been Second, effect sizes were calculated using from tumor volume of effect sizes may have but were between Third, experimental design apparent in 'Materials and methods' our failure to detect a dose–response For instance, few reports provide animal and testing to important in tumor growth. It should also be that our study was in findings like will to be using our study analysis of a single and it may be our findings do not receptor tyrosine kinase or sunitinib. However, many of our findings are consistent with those observed in other systematic of preclinical cancer (Amarasingh et al., 2009; et al., Hirst et al., 2013). our analysis not address many design duration of experiment or of are to on study validity. Finally, we that there may be funding that limit of validity practices described We that other in particular, have found to make such methods a commentators have concerns about the design and reporting of preclinical cancer research et al., Begley, In one report, only 11% preclinical cancer studies to a major replication (Begley and Ellis, 2012). The for Open and has a that will to of the highest impact in cancer published between and 2014). In a recent et al. many researchers for in preclinical using drug that are due to toxicity and Houghton, 2013). preclinical validity threats like the in our clinical development trajectories. Many research like and have design guidelines at improving the clinical generalizability of preclinical studies et al., 2009; et al., et al., et al., and the guidelines (Kilkenny et al., for reporting animal experiments have been taken up by numerous and funding Our findings provide further for and guidelines for the design, and of preclinical studies in cancer. Materials and methods a identify all in animal studies testing the anti-cancer of sunitinib we the on February using a from et al. and et al. and of coverage from to and database of coverage from to and of coverage from to 2012). Search results were into an and were were identified the of identified articles. Table 1—source data for and PRISMA was performed at level by two and and at by one were original reports or English at least one experiment response in a non-human animals, and employed sunitinib in a or experimental tested anti-cancer activity. the same experiment in where the same experiment was reported in different the most recent publication was a All included studies were evaluated at the but only those with eligible experiments (e.g., those the effect of monotherapy on tumor volume and that were reported with sample sizes and error were to We excluded experiments when they had been reported in a publication after for and For each eligible we extracted experimental design elements from a systematic review of validity threats in preclinical research (Henderson et al., 2013). the of internal and construct validity are given in Figure 1—source data for external validity, we an that the number of species and models tested for a given malignancy and an extra point if more than one species and model was For example, if experiments within a malignancy tested two species and three different model the external validity would be 4 point for the second one point for the second model one point for the model and an extra point because more than one model and species were Our outcome was experimental tumor volume and we extracted information mean of treatment effect, and to of study and level effect sizes. the of tumor volume were not consistent between experiments, we extracted those experiments for which a of tumor volume could be These included reported in or tumor reported in from tumor cell lines reported in and in tumor between the control and treatment arms. We extracted experiments of both and but not experiments where tumor was reported. for these different measures of tumor SMDs were calculated using Hedges' Hedges' is a widely standardized of effect in where are not For experiments where more than one dose of sunitinib was tested against the same control we a pooled SMD to for the use of the same control were extracted at and as the first of drug 14 measured data point to 14 days first and the last common time point between the control group and the treatment The was between experiments and the last time point for which we could calculate SMD and the point at which the greatest difference was observed between the arms. were extracted using the was performed by and and using was a to heterogeneity and prevent in were and if by a The rate before for all studies was a Effect sizes were calculated as SMDs using Hedges' with 95% confidence Pooled effect sizes were calculated using a random effects model employing the and in (Wallace et al., We also calculated heterogeneity within each malignancy using statistics (Figure 2—source data the predictive value of preclinical studies in our sample, we calculated pooled effect sizes for each type of analyses were performed for each validity were calculate
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Étiquettes directes de modèles (non validées)
Étiquettes de catégorie et de devis d'étude par modèle, issues des rondes d'étiquetage. C'est une sortie machine, non validée, et le désaccord entre modèles est livré comme donnée. Aucun devis ici n'est encore validé contre MEDLINE.
| Bras | Catégories | Devis d'étude | Confiance |
|---|---|---|---|
| gpt | Métarecherche Domaine: Méthodes · Genre: Commentaire Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non | Sans objet | low |
| grok | aucune catégorie Domaine: non disponible · Genre: Commentaire Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non | Sans objet | low |
| opus | MétarechercheMéta-épidémiologie (sens large) Domaine: Méthodes · Genre: Autre Porte sur le système de recherche canadien: non · Porte sur un sujet canadien: non | Sans objet | low |
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,057 | 0,105 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,006 | 0,003 |
| Bibliométrie | 0,002 | 0,003 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,001 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéeÉtiqueté directement par 3 modèles lisant le dossier complet.
Les modèles divergent sur des parties de cette classification; chaque voix est préservée dans la section en fin de page.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».