Methodological concerns and clinical relevance: a critical appraisal of the meta-analysis on biosimilars versus originator follitropin alfa in ART
Bibliographic record
Abstract
Dear Editors, We read with interest the systematic review and meta-analysis by Kiose et al. (2025), which reports that rates of live birth, clinical pregnancy, and ongoing pregnancy are significantly lower following treatment with follitropin alfa biosimilars than with the originator product, Gonal-f®, in women undergoing ovarian stimulation (OS) for ART. The results are reported with moderate-to-low certainty of evidence, yet the authors conclude that clinicians should be informed ‘pregnancy rates after a fresh transfer are likely to be lower’ following OS with biosimilars than with originator follitropin alfa. Overall, the meta-analysis included eight randomized controlled trials (RCTs), and pooled data from 2987 women and seven different ‘biosimilars’ versus the originator. A fundamental limitation of the meta-analysis, which uses a standard pairwise design to compare two interventions, is the misrepresentation of ‘biosimilar’ follitropins approved by different regulators as a homogeneous comparator group, without distinguishing potential variability in molecular characteristics or pharmacokinetics/pharmacodynamics. Inherent molecular variability is expected among complex biological medicines, even among batches of the originator product. Strict and aligned criteria for registering biosimilars have been adopted by WHO Listed Authorities (WLAs) including Europe, Canada, the USA, Australia, and Japan, among which molecular differences must be contained within an accepted variability that is not expected to affect therapeutic equivalence (de Mora and Fauser, 2017). Comparable pharmacological, pharmacokinetic, efficacy, and safety profiles must also be demonstrated. Other regulators around the world have approved follitropins under less stringent criteria. Indeed, detailed analysis of originator follitropin alfa versus products approved in regions outside Europe has demonstrated several differences in structural features believed to affect biological activity (Manzi et al., 2022). The meta-analysis combines data from two WLA-approved biosimilars (Bemfola®/Afolia® and Ovaleap®) and five products for which biosimilarity has not been proven according to the rigorous WLA regulatory standards (Primapur®, QL1012, Cinnal-f®, Folitime®, Follitrope®). Although subgroup analysis on the primary outcome measure reportedly found no significant difference between subgroups when studies were grouped by Afolia®/Bemfola®/Ovaleap® or others, the meaning of this result is difficult to interpret as no further details were presented, and the small number of studies included is too few to allow meaningful conclusions on heterogeneity to be made. Other aspects of the meta-analysis design also require scrutiny. The validity of live birth rate (LBR) as a primary outcome is questionable given that only three of the studies included LBR among planned outcome measures, and none were powered to demonstrate differences in LBR. Most studies measured the number of oocytes retrieved as the primary endpoint and were powered for non-inferiority of test product versus originator. Importantly, the meta-analysis showed no significant differences in oocyte numbers, or in rates of ovarian hyperstimulation syndrome. OS with follitropin alfa is only one step in the complex ART process and potential variability in other procedural and patient factors that impact implantation success, pregnancy outcomes, and live birth were not considered in the analysis. Furthermore, seven of the eight studies were only single-blinded, potentially influencing patient-related factors. There are additional concerns about the quality and comparability of studies pooled in the analysis. Three of the eight RCTs were judged as having overall high risk of bias due to deviations from intended intervention and missing outcomes data, and a further two were of concern due to potential randomization bias or missing outcomes data. Some of these studies were excluded in sensitivity analyses, but the remaining analysis populations were small. All studies were industry-sponsored. The analysis reported no statistical heterogeneity among the RCTs. However, descriptive data reveal differences in follitropin alfa dosing schedules, criteria for hCG administration, study populations, and baseline patient characteristics, none of which was explored in sensitivity analyses. The largest of the studies (N = 1101 US women; Fertility Biotech AG, 2017), which was not published in a peer-review journal nor included in regulatory submissions, represents over one-third of the total analysis population, and was conducted in older women (35–42 years), whereas other studies enrolled younger populations of 18–20 to 35–39 years of age. This study was not separated in sensitivity analyses, despite additional concerns for risk of bias arising from missing randomization information. Finally, the exclusion of poor responders and other populations of interest in most of the studies limits generalizability of the results to real-world ART settings, where cost-effectiveness and convenience are also important considerations. Real-world evidence from 17 ART centers throughout France has shown that cumulative LBR per stimulated cycle with biosimilar follitropin alfa is not significantly different than originator follitropin alfa (Barrière et al., 2023) and is more cost-effective from a French healthcare-system perspective (Lehmann et al., 2024). Overall, we propose that the meta-analysis is fundamentally flawed by poor statistical methodology, lack of comparability among pooled comparators and studies, and questionable quality of some of the study inputs. The conclusions should be interpreted with great caution and do not provide robust evidence to deter patients and clinicians from considering rigorously approved biosimilar products equally among their options for follitropin alfa OS during ART. P.S. has received research grants for his institution from Besins, Ferring, Gedeon Richter, Merck, and Theramex; speaker fees/honoraria and support for attending meetings from Besins, Ferring, Gedeon Richter, IBSA, Merck, and Theramex; and serves as a board member of the Society of Endometriosis and Uterine Disorders, and editorial board member of Reproductive BioMedicine Online and Gynécologie Obstétrique Fertilité & Sénologie. M.I.L. has received support for attending meetings from CER, Gedeon Richter, and MSD Laboratory Chile. R.M. has received research support from Theramex; speaker fees/honoraria from Gedeon Richter; and is a member of the president’s council of the Italian Society of Human Reproduction. T.F. has received speaker fees and support for meeting attendance from Gedeon Richter. P.B., G.D., M.G., S.H., and B.S. have no conflicts of interest to declare.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.199 | 0.500 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.010 | 0.014 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.001 | 0.004 |
| Scholarly communication | 0.007 | 0.004 |
| Open science | 0.005 | 0.002 |
| Research integrity | 0.013 | 0.013 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".