Intravenous immunoglobulin and idiopathic secondary recurrent miscarriage: methodological problems
Bibliographic record
Abstract
Sir, Stephenson et al. (2010) report a double-blind randomized controlled trial (RCT) of intravenous immunoglobulin (IVIG) in idiopathic secondary recurrent miscarriage and conclude ‘no benefit was found’. A variety of serious methodologic flaws question the integrity of this conclusion. First, the planned sample size was not achieved due to slow recruitment, and the authors chose to make a report based on an underpowered study, a well-known error (Altman and Bland, 1995; Halpern et al., 2002). Meta-analysis gave a small positive benefit in the patients receiving IVIG, but Stephenson et al. failed to consider how large the trial would have had to have been to confidently conclude that this represented ‘no treatment effect’. The number may be estimated at 407 in each group (Clark, 2011); one calculates a Pα value at the end of a study, to avoid a type I error, and there is no reason why one should not also calculate a Pβ error, to estimate the risk of a type II error, and to determine how large the sample should have been for adequate power given the true (observed) treatment effect size was less than the apparently over-optimistic guesstimate at the beginning when the study was designed. Hence, the correct conclusion for the Stephenson et al. study should have been that no conclusion is possible given the inadequate sample size achieved. Second, the authors set out to study an immunomodulatory treatment but did not think it wise to determine if the subjects had any detectable immune abnormalities that would merit their inclusion in a study of a putative immune therapy. Indeed, they excluded from their meta-analysis two trials by Christiansen et al. (1995, 2002) because some of the patients in these studies had immune abnormalities. One would not logically expect a treatment to work in a patient who did not have a condition that the treatment could correct, so what physiologic abnormality did Stephenson et al. think they might be correcting with IVIG treatment? If the Christiansen et al. data are included in the meta-analysis, the net result is a statistically significant positive benefit (Fig. 1A). Subdividing the trials as includable or not includable in a meta-analysis, as done by Stephenson et al., shows a rather striking difference (Fig. 1B); the success rates with IVIG treatment were similar in the two groups, but the control group in the Christiansen et al. trials had a much lower spontaneous success rate. It is much easier to detect a positive treatment effect when the patients have a condition causing a high rate of loss, and a simple power analysis shows the total sample size in the Christiansen et al. RCTs was more than adequate for 80% power (Clark, 2011). Small studies with inadequate power, in contrast, produce results which are untrustworthy, whether the Pα value is <0.05 or >0.05 (Clark et al., 2001). (A) Meta-analysis of randomized controlled trials of IVIG in secondary recurrent miscarriage (Peto method). Odds ratio was 2.10 (95% confidence interval 1.06–4.49) in favor of IVIG treatment, 2P = 0.034; χ2 = 2.58 (no significant heterogeneity precluding combination of the five studies). (B) Comparison of control and IVIG-treated patients that were included as ‘idiopathic’ by Stephenson et al. (2010) to those excluded. *Significantly lower than success rate in the ‘idiopathic’ included groups (P< 0.001); **significantly better than control (P< 0.01) but not different from the ‘idiopathic’ IVIG-treated group (P > 0.2). The number per group to achieve a power of 80% for the difference shown was 407 for the ‘idiopathic’ included group and 22 for the excluded group. Adapted from the American Journal of Reproductive Immunology (Clark, 2011). Third, whilst CD56+16− NK cells that predominate in the uterus may play a role in enhancing neovascularization of the implanting embryo, it is CD56+16+ type NK cells that predominate in circulating blood that are considered important in the pathogenesis of recurrent miscarriage (Clark, 2008). Contrary to citation of Lachapelle et al. (1996), Clifford et al. (1999) and Quenby et al. (1999) as showing a role for CD56+16− cells, Lachapelle et al. and Quenby et al. showed an excess of CD56+16+ cells in the preimplantation endometrium of women with recurrent miscarriage and Clifford et al. studied only CD56+ cells without reference to their CD16 phenotype. It has been reported by Yamada et al. (2003) that women aborting normal karyotype embryos had elevated blood CD56+16+ cells in contrast to women aborting abnormal embryos where blood NK cell levels were normal. Not all IVIG preparations are equally effective in suppressing elevated blood-type NK cells, and notably, Gamimune which was used for some patient in the current Stephenson et al. study may be less potent than other types of IVIG (Clark et al., 2008). If a treatment is suboptimal, so would be the clinical outcome, and the only way to be sure a treatment is correcting a physiologic abnormality is to measure it! Lack of attention to immune abnormalities, the immunobiology of treatment regimens, and the inability of a small meta-analysis to achieve the required power can lead to unsupportable beliefs. It is fortunate that Fig. 1 suggests a remedy with respect to secondary recurrent miscarriage [It should be noted that a recent observational study of IVIG in IVF failure patients has shown a similar result: only patients with immune test abnormalities showed a statistically significant treatment benefit with the sample size available (Winger et al., 2011)]. It would be more profitable in the case of so-called idiopathic recurrent miscarriage to work out mechanisms before expending large amounts of time, money and energy on conducting inconclusive RCTs in the hope that a convincing result can be obtained. With respect to the effects of paternal mononuclear cell treatment in unexplained primary recurrent miscarriage, about which Stephenson et al. suggest inefficacy, a more up-to-date appreciation of the literature on the biology and statistical considerations may alter antiquated notions (Clark, 2008). With respect to IVIG in secondary recurrent miscarriages, as with other syndromes of recurrent pregnancy loss, one may profitably begin by attempting to determine which test abnormalities and which treatment effects may correlate with a normal live birth in the control group and in the IVIG group (Yamada et al., 2003; van den Heuvel et al., 2007; Winger et al., 2011).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.107 | 0.478 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.002 | 0.006 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.019 | 0.014 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".