Need for equivalence testing of efficacy of alternative antibiotics for treatment of pertussis
Bibliographic record
Abstract
To The Editors: The recommended antibiotic for treatment and prophylaxis of pertussis in the United States and Canada is 14 days of erythromycin treatment. Because erythromycin is associated with adverse gastrointestinal effects, a 7-day regimen of clarithromycin was evaluated for efficacy and safety in a recent clinical trial. 1 The results of this trial provide valuable information about the efficacy and safety of clarithromycin. Given the importance of the results for public health and medical practice, we would like to point out two limitations of the study that should be considered in evaluation of the results. The results did not include the appropriate statistical analysis for comparing the efficacy and safety of the two antibiotics, and the sample size of the clinical trial was insufficient to support the conclusion that clarithromycin and erythromycin were equally effective agents. When we performed equivalence testing on the data collected to determine efficacy, our results suggested that clarithromycin might be less efficacious than erythromycin. Equivalence testing is the appropriate analysis of efficacy because the design of the study was an equivalence trial 2 to evaluate whether clarithromycin was similar to the “standard” treatment with erythromycin. Instead of equivalence testing, the analysis of efficacy data for the per protocol treatment groups consisted of calculating microbiologic eradication rates and Clopper-Pearson (exact binomial) 3 95% confidence intervals (CIs): 100% (31 of 31), 95% CI 88.8 to 100 for clarithromycin vs. 95.7% (22 of 23); 95% CI 78.1 to 99.9 for erythromycin. 1 The interpretation that these results demonstrated equivalence may have been erroneously based on the observation that the confidence intervals for the respective microbiologic eradication rates overlapped, although this observation was not explicitly stated in the paper. In any case examining the overlap between confidence intervals to judge the statistical significance of differences should not be used for formal significance testing. 4 Equivalence testing is used in studies intended to show the equivalence of two rates; i.e., the two rates do not differ by more than a specified amount. This is in contrast to testing the hypothesis of no difference between the rates, which is intended to show only that the two rates are different. In equivalence testing, a difference in rates less than a prespecified quantity is considered to be not important. A commonly used equivalence test is the two one-sided test (TOST) procedure, which tests two one-sided hypotheses simultaneously. 5 We applied the TOST procedure to the efficacy data presented by Lebel and Mehra 1 to test the null hypothesis that the absolute value of the difference between the microbiologic eradication rates for clarithromycin and erythromycin was greater than or equal to a prespecified quantity vs. the alternative hypothesis that it was less than the prespecified quantity. We used a prespecified quantity of 10% because the sample size per treatment group in the trial was determined by assuming a maximum difference between treatments of 10%. 1 Assuming that the level of significance is 5% for each one-sided test, the TOST procedure rejects the null hypothesis and concludes that the rates are equivalent if a 90% two-sided CI for the observed difference in rates is completely contained in the equivalence interval, (−10% to +10%). Confidence limits for the observed difference in microbiologic eradication rates can be computed using asymptotic (large sample) or exact (small sample) methods. 6 Our analysis using the 5% level TOST procedure and the asymptotic method for calculating confidence limits failed to reject the null hypothesis because the asymptotic 90% CI for the observed difference in microbiologic eradication rates between clarithromycin and erythromycin, (−2.6% to 11.3%), was not completely contained in the equivalence interval, (−10% to +10%) (Table 1). This result suggested that clarithromycin was more efficacious than erythromycin. However, given the small sample sizes in each treatment group, a more prudent approach to the analysis would rely on the exact method of calculating the confidence limits. The results of the exact method also failed to prove equivalence of the two rates. For instance the exact 90% CI for the difference in rates was (−12.1% to 29.8%), which suggested that the rate for clarithromycin could be as much as 12% worse than that of erythromycin. The exact upper 90% confidence limit suggested that the rate for clarithromycin may be up to 30% higher than that of erythromycin.TABLE 1: Statistics used in equivalence testing* for evaluating the efficacy of clarithromycin and erythromycin for the per protocol treatment groupsThe other limitation of the study was the number of patients available to assess efficacy. The authors noted that a sample size of 126 subjects (63 subjects per treatment group) was required to detect a maximum difference between treatments of 10% with a probability of 90% (α = 0.10, one-sided). 1 We could not replicate this sample size calculation, but we are not certain which method of sample size determination was used. 7 Nonetheless the authors assumed that 50% of the cultures would be negative for Bordetella pertussis, and the dropout rate would be 15%, such that the total sample size required was 300 subjects (150 subjects per treatment group). Patient enrollment was ended prematurely; only 51% (153) of the targeted 300 subjects were enrolled and randomized. The final sample size in each per protocol treatment group was small (31 patients for clarithromycin and 23 patients for erythromycin) and lowered the statistical power to detect a difference in the microbiologic eradication rates. The sample size required for equivalence testing of the microbiologic eradication rates can also be calculated under statistical assumptions similar to those used in the study. Assuming that the expected microbiologic eradication rate for both erythromycin and clarithromycin was 95%, to detect a maximum difference of 10% between the rates, with a significance level of 10% (two-sided) and a statistical power of 90%, the required sample size was 206 subjects (103 subjects per treatment group). 8 Thus not enough subjects were enrolled for either testing the hypothesis that the two microbiologic eradication rates are equivalent or testing the hypothesis that the two rates are equal. In summary, equivalence testing of microbiologic eradication rates is the appropriate analysis for evaluating the efficacy of alternative antibiotic treatments in equivalence trials. Equivalence testing should be conducted for all of the outcomes in equivalence trials, including clinical cure rates and rates of adverse events and compliance. Our results of equivalence testing of efficacy data from the recent trial 1 comparing the efficacy of clarithromycin vs. erythromycin warrant a limited recommendation for use of clarithromycin as an alternative choice for treatment of pertussis. Additional studies are needed to determine the effectiveness of alternative antibiotics for prophylaxis and treatment of pertussis. Our analysis of the recent efficacy data for clarithromycin highlights the importance of considering the study objective and accounting for the study design in statistical analysis. Andrew L. Baughman, Ph.D., M.P.H. Kristine M. Bisgard, D.V.M., M.P.H.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.158 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".