MétaCan
Menu
Back to cohort

Need for equivalence testing of efficacy of alternative antibiotics for treatment of pertussis

2003· letter· en· W2050518819 on OpenAlexaboutno aff
Andrew L. Baughman, Kristine M. Bisgard

Bibliographic record

VenueThe Pediatric Infectious Disease Journal · 2003
Typeletter
Languageen
FieldMathematics
TopicStatistical Methods in Clinical Trials
Canadian institutionsnot available
Fundersnot available
KeywordsClarithromycinErythromycinMedicineRegimenClinical trialAdverse effectRandomized controlled trialInternal medicineConfidence intervalAntibioticsEquivalence (formal languages)MathematicsBiologyHelicobacter pyloriMicrobiology

Abstract

fetched live from OpenAlex

To The Editors: The recommended antibiotic for treatment and prophylaxis of pertussis in the United States and Canada is 14 days of erythromycin treatment. Because erythromycin is associated with adverse gastrointestinal effects, a 7-day regimen of clarithromycin was evaluated for efficacy and safety in a recent clinical trial. 1 The results of this trial provide valuable information about the efficacy and safety of clarithromycin. Given the importance of the results for public health and medical practice, we would like to point out two limitations of the study that should be considered in evaluation of the results. The results did not include the appropriate statistical analysis for comparing the efficacy and safety of the two antibiotics, and the sample size of the clinical trial was insufficient to support the conclusion that clarithromycin and erythromycin were equally effective agents. When we performed equivalence testing on the data collected to determine efficacy, our results suggested that clarithromycin might be less efficacious than erythromycin. Equivalence testing is the appropriate analysis of efficacy because the design of the study was an equivalence trial 2 to evaluate whether clarithromycin was similar to the “standard” treatment with erythromycin. Instead of equivalence testing, the analysis of efficacy data for the per protocol treatment groups consisted of calculating microbiologic eradication rates and Clopper-Pearson (exact binomial) 3 95% confidence intervals (CIs): 100% (31 of 31), 95% CI 88.8 to 100 for clarithromycin vs. 95.7% (22 of 23); 95% CI 78.1 to 99.9 for erythromycin. 1 The interpretation that these results demonstrated equivalence may have been erroneously based on the observation that the confidence intervals for the respective microbiologic eradication rates overlapped, although this observation was not explicitly stated in the paper. In any case examining the overlap between confidence intervals to judge the statistical significance of differences should not be used for formal significance testing. 4 Equivalence testing is used in studies intended to show the equivalence of two rates; i.e., the two rates do not differ by more than a specified amount. This is in contrast to testing the hypothesis of no difference between the rates, which is intended to show only that the two rates are different. In equivalence testing, a difference in rates less than a prespecified quantity is considered to be not important. A commonly used equivalence test is the two one-sided test (TOST) procedure, which tests two one-sided hypotheses simultaneously. 5 We applied the TOST procedure to the efficacy data presented by Lebel and Mehra 1 to test the null hypothesis that the absolute value of the difference between the microbiologic eradication rates for clarithromycin and erythromycin was greater than or equal to a prespecified quantity vs. the alternative hypothesis that it was less than the prespecified quantity. We used a prespecified quantity of 10% because the sample size per treatment group in the trial was determined by assuming a maximum difference between treatments of 10%. 1 Assuming that the level of significance is 5% for each one-sided test, the TOST procedure rejects the null hypothesis and concludes that the rates are equivalent if a 90% two-sided CI for the observed difference in rates is completely contained in the equivalence interval, (−10% to +10%). Confidence limits for the observed difference in microbiologic eradication rates can be computed using asymptotic (large sample) or exact (small sample) methods. 6 Our analysis using the 5% level TOST procedure and the asymptotic method for calculating confidence limits failed to reject the null hypothesis because the asymptotic 90% CI for the observed difference in microbiologic eradication rates between clarithromycin and erythromycin, (−2.6% to 11.3%), was not completely contained in the equivalence interval, (−10% to +10%) (Table 1). This result suggested that clarithromycin was more efficacious than erythromycin. However, given the small sample sizes in each treatment group, a more prudent approach to the analysis would rely on the exact method of calculating the confidence limits. The results of the exact method also failed to prove equivalence of the two rates. For instance the exact 90% CI for the difference in rates was (−12.1% to 29.8%), which suggested that the rate for clarithromycin could be as much as 12% worse than that of erythromycin. The exact upper 90% confidence limit suggested that the rate for clarithromycin may be up to 30% higher than that of erythromycin.TABLE 1: Statistics used in equivalence testing* for evaluating the efficacy of clarithromycin and erythromycin for the per protocol treatment groupsThe other limitation of the study was the number of patients available to assess efficacy. The authors noted that a sample size of 126 subjects (63 subjects per treatment group) was required to detect a maximum difference between treatments of 10% with a probability of 90% (α = 0.10, one-sided). 1 We could not replicate this sample size calculation, but we are not certain which method of sample size determination was used. 7 Nonetheless the authors assumed that 50% of the cultures would be negative for Bordetella pertussis, and the dropout rate would be 15%, such that the total sample size required was 300 subjects (150 subjects per treatment group). Patient enrollment was ended prematurely; only 51% (153) of the targeted 300 subjects were enrolled and randomized. The final sample size in each per protocol treatment group was small (31 patients for clarithromycin and 23 patients for erythromycin) and lowered the statistical power to detect a difference in the microbiologic eradication rates. The sample size required for equivalence testing of the microbiologic eradication rates can also be calculated under statistical assumptions similar to those used in the study. Assuming that the expected microbiologic eradication rate for both erythromycin and clarithromycin was 95%, to detect a maximum difference of 10% between the rates, with a significance level of 10% (two-sided) and a statistical power of 90%, the required sample size was 206 subjects (103 subjects per treatment group). 8 Thus not enough subjects were enrolled for either testing the hypothesis that the two microbiologic eradication rates are equivalent or testing the hypothesis that the two rates are equal. In summary, equivalence testing of microbiologic eradication rates is the appropriate analysis for evaluating the efficacy of alternative antibiotic treatments in equivalence trials. Equivalence testing should be conducted for all of the outcomes in equivalence trials, including clinical cure rates and rates of adverse events and compliance. Our results of equivalence testing of efficacy data from the recent trial 1 comparing the efficacy of clarithromycin vs. erythromycin warrant a limited recommendation for use of clarithromycin as an alternative choice for treatment of pertussis. Additional studies are needed to determine the effectiveness of alternative antibiotics for prophylaxis and treatment of pertussis. Our analysis of the recent efficacy data for clarithromycin highlights the importance of considering the study objective and accounting for the study design in statistical analysis. Andrew L. Baughman, Ph.D., M.P.H. Kristine M. Bisgard, D.V.M., M.P.H.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.158
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.947
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0020.158
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.408
GPT teacher head0.499
Teacher spread0.090 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations4
Published2003
Admission routes1
Has abstractyes

Explore more

Same venueThe Pediatric Infectious Disease JournalSame topicStatistical Methods in Clinical TrialsFrench-language works237,207