The use of the noninferiority analysis in clinical studies
Bibliographic record
Abstract
This issue of the Equine Veterinary Journal presents a number of study reports that examined whether a type of medical treatment was better than another one that differed by the timing of drug administration 1, the administered dose level 2, or the type of drug administered 3, 4. This type of information is extremely appealing to practitioners, as it may provide explicit recommendations with respect to the choice of drug or dosing regimen in the clinical setting. However, the research hypothesis being tested in these studies differs from the usual one of a difference between treatments, and is a noninferiority (NI) analysis, which is essentially a one-sided equivalence analysis. The NI analysis could be used to determine if the alternative treatment would be acceptable. The statistical tests in 3 of the studies 2-4 determined whether 2 treatments did not significantly differ from each other. In contrast, Sykes et al. 2 used the NI approach in compliance with the stated hypothesis. The purpose of this editorial is to introduce the NI analysis, and to present some of the advantages and disadvantages of its use as a statistical tool. In the drug approval process, a NI approach can be used to compare an investigational drug with an approved drug, which acts as the active control, to determine if the investigational drug is noninferior to the active control. The NI analysis allows the investigational drug to be less effective than the active control by a predetermined amount and still be considered effective. The use of the NI evaluation is described in the Food and Drug Administration's (FDA) Center for Veterinary Medicine Guidance for Industry (GFI) #204 Active Controls in Studies to Demonstrate Effectiveness of a New Animal Drug for Use in Companion Animals 5. This guidance represents FDA's current thinking and recommendations, and is not a legally binding document. This guidance is a useful resource when considering the use of a NI approach because it describes both the NI evaluation and the margin of difference (delta), and provides general considerations when conducting a NI study. In this editorial, we will extend these concepts to the nonregulatory context in order to include the testing of different drugs, dosing schedules and dosing conditions outside the drug approval process. Another useful resource is Friese et al.'s publication 6, which is a comprehensive description of the use of the NI analysis in veterinary clinical trials. The study objective in many clinical studies is to compare the effectiveness (and/or safety) of 2 treatments with the collected data and arrive at a clinically meaningful verification of the study hypotheses. The type of hypothesis depends on whether the investigator is interested in establishing that the treatments provide similar responses, or that one of the treatments is superior to the other. In comparison with a superiority study, the null hypothesis in the NI study is that the difference between the 2 treatments is greater than the proposed maximum acceptable difference between treatment outcomes (delta). The study is designed to reject the null hypothesis and conclude that the difference between the 2 treatments is less than delta 7. In some cases, the NI null hypothesis may be more reasonable to propose than the usual superiority analysis, because it is based on prior effectiveness or safety information. For instance, de Lagarde et al. 4 could justify their hypothesis by stating that the positively charged quaternary ammonium of N-butyl-scopolammonium plays a critical role in limiting its diffusion to the central nervous system, and therefore the drug is expected to produce fewer adverse events 8. A 2-sided 95% confidence interval (CI) for the difference between the animals treated with the active control vs. the investigational drug is calculated. If the upper bound of this CI is less than the margin of difference, then the investigational drug is considered to be noninferior to the active control. To illustrate the concept of NI, GFI 204 provides Figure 1 depicting five hypothetical NI scenarios 5. The studies were analysed by calculating from the hypothetical study data a 2-sided 95% CI for the difference in per cent cures, active control minus investigational new animal drug. This difference is represented by the diamond on each line on the graph. The right end of the line depicts the upper confidence bound. If the upper bound of the CI is <15%, then one can conclude that the investigational new animal drug is noninferior. That is, if the right side of the horizontal line does not cross the vertical line at 15% (defined margin of difference) there is sufficient evidence to conclude the investigational new animal drug is noninferior to the active control. If the right side of the horizontal line crosses the vertical line at 15% (defined margin of difference) there is insufficient evidence to conclude the investigational new animal drug is noninferior to the active control. In this example, Lines 2, 4 and 5 can be said to have sufficient evidence to conclude that the investigational new animal drug is noninferior to the active control. More detail can be found in GFI 204 5. Five hypothetical noninferiority scenarios with an active control. (Margin of difference = 15%; randomisation = 1:1; α = 0.05.) There are some limitations to the interpretation of a P value in statistical analysis. To choose a null hypothesis, and then to arrive at a P value that is used to reject or fail to reject the null hypothesis, may not provide a relevant conclusion, as the difference between the 2 treatment groups may be statistically significant but clinically meaningless. Furthermore, a P value does not provide information about the size or direction of the effect or the range of possible outcomes. In contrast, a CI based on the difference between 2 treatments provides a more comprehensive measure of the difference between treatment effects. The 95% CI can be defined as follows: if the experiment were repeated 100 times, 95 out of the 100 ranges collected would contain the true parameter that one was measuring. However, for a particular 95% CI the range either contains the true parameter or not. The CI not only evaluates the null hypothesis but indicates the effect size (the estimated relative effects of the 2 treatments) and the lower and upper bounds of the estimated difference 7. These bounds can be used to evaluate superiority of one treatment vs. another using a superiority analysis, or to demonstrate that one treatment is not inferior to the other using the NI analysis. Therefore, the use of the CI in the analysis of comprehensive clinical endpoints that are the result of multiple variables, such as response to treatment at different time points and clinical pathology results, will provide more inferential value to the study's results. The decision to use the NI analysis for a study warrants some specific choices in study design, data analysis and interpretation of the study results. A critical but often overlooked feature of NI studies is that the study design must provide a fair challenge between the treatments to be compared, and both must be relevant therapeutic options. If this is not the case, a drug may appear inferior to its comparator so the investigators erroneously infer that it is worse, when in fact it was used with inadequate consideration to its pharmacokinetic or pharmacodynamic properties. Some examples of this situation are: 1) a challenge infection with a pathogen whose level of antimicrobial sensitivity significantly differs between the tested antibiotics; 2) using a dose size or dosing interval that allowed one drug to accumulate and reach its effective (or toxic) plasma concentration but not the comparator drug; and 3) measuring an endpoint at a time that allowed the total elimination of one drug but not the comparator drug. Investigators may prevent this problem by performing trial simulations using a pharmacokinetic–pharmacodynamic model that may be combined with a disease progression model if the study is performed in the early stages of disease 9. For example, the decision to administer meloxicam and flunixin meglumine every 12 h instead of every 24 h in the study by Naylor et al. 3 was supported by pharmacokinetics and clinical experience. In an NI analysis, an acceptable margin of difference (delta) between the 2 treatments is established a priori. Among the studies referenced previously 1-4, a delta was established a priori in one study of Sykes et al. 2. In their other study 1, Sykes et al. chose a delta of ≥20% difference in ulcer healing as an acceptable difference in treatment response to perform post hoc power calculations. The same approach to choosing an acceptable difference could have been used to establish the delta a priori. The delta is very specific to the active control, and is generally based on the expected difference between treatment outcomes (effectiveness) for the active control vs. a placebo (or other comparator), and the maintenance of a proportion of the effectiveness by the investigational drug. For example, if it has been demonstrated that the active control results in a 20% improvement in treatment success compared with placebo, the margin might be set at 10% of this value, thus preserving at least 50% of the effectiveness. The calculation of the margin of difference is the most complex and controversial aspect of the NI; it is based on the effect size of the active control, which in most cases in the drug approval process was demonstrated in a placebo-controlled effectiveness study. In the absence of placebo-controlled data, the effect size of the active control could be estimated from historical evidence, such as scientific literature, expert opinion, or nonplacebo controlled studies. One of the underlying assumptions of NI is that the active control will have the same effectiveness in the NI study as it had in the original study, so a conservative active control effect size is chosen based on the difference between the active control and the comparator (e.g. placebo control) in the original study. The margin of difference to be used in the assessment of NI between the investigational drug and active control is generally some portion of this effective size. In general, the margin of difference should conserve at least 50% of the effect size. For example, if the effect size between the placebo and active control was 20%, then the margin of difference between the investigational drug and active control would be no more than 10%. The assay sensitivity of the NI study is measured by its ability to distinguish an effective treatment from a less effective or ineffective treatment. The sensitivity depends on the active control maintaining the same level of effectiveness as it had in the original study. The sensitivity of the study is further enhanced if the study design uses similar outcome variables, such as enrolment, treatment and evaluations, to those used to determine the effectiveness of the active control in the original study. The NI approach, given that the assumptions are met, provides assurance that the difference between the 2 treatments is small enough to be clinically acceptable. A statistical test for superiority, by contrast, may conclude that the treatments are significantly different but can make no inference regarding the clinical relevance of the difference between the 2 groups. The NI study approach can be designed to incorporate multiple types of clinical endpoints, such as dichotomous scales (success/failure), ordinal scales (mild, moderate, severe), and continuous variables, such as clinical pathology values. Depending on the treatments and indication, an NI study may be more ethical than a placebo-controlled study, because all animals receive treatment. Consequently, NI studies may have higher enrolment and retention of client-owned animals, whose owners do not want to take the risk that their animal will be in the placebo group. If the investigational drug is more effective than the approved drug, an NI study may use fewer animals than a superiority study 7. The objective of the NI test may not be as clear as that of the superiority test, because in comparison, the null and alternative hypotheses are essentially reversed. The superiority study is designed to reject the null hypothesis of equal treatment groups in favour of the alternative hypothesis that the treatment groups are different. The NI study is designed to reject the null hypothesis that the investigational drug is worse than the active control by a clinically relevant margin in favour of the alternative hypothesis that the investigational drug is not clinically worse than the active control. There is also a risk of falsely concluding NI if the assumption that the active control was as effective as in the placebo-controlled study is false. In theory, this could result in a study in which the investigational drug is proven to be noninferior to the active control, but both treatments are ineffective. The margin of difference is often criticised because it is based on the assumption that the active control will perform exactly as it did in the placebo-controlled study. In addition, inherent variability among studies, such as patient characteristics, concomitant medications, owner compliance, and temporal influences, also influence the validity of the study 10. To compensate for this variability, the NI study should be designed to replicate the placebo-controlled study as closely as possible. The number of animals needed for adequate statistical power in a NI study is influenced by the margin of difference and the estimated outcome differences for both the investigational drug and active control. In many cases, a NI study may require more animals for adequate statistical power than a placebo-controlled study. In general, more animals would be needed for a NI study with a high self-cure rate for the disease, a small margin of difference, and a lower cure rate for the investigational drug compared with the active control. In conclusion, the NI analysis should be considered when the study objective is to compare 2 treatments within a clinically acceptable margin. There are some disadvantages to this type of analysis, such as the reliance on the previous performance of the active control, and the complexity of determining the margin of difference, but these may be outweighed by its advantages, such as the incorporation of multiple types of endpoints, and the use of an active control (vs. a placebo control), which addresses ethical concerns and may increase enrolment and retention of client-owned animals.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.040 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.002 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.004 | 0.004 |
| Research integrity | 0.001 | 0.008 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".