Applying the win ratio method in clinical trials of orphan drugs: an analysis of data from the COMET trial of avalglucosidase alfa in patients with late-onset Pompe disease
Bibliographic record
Abstract
BACKGROUND: Clinical trials for rare diseases often include multiple endpoints that capture the effects of treatment on different disease domains. In many rare diseases, the primary endpoint is not standardized across trials. The win ratio approach was designed to analyze multiple endpoints of interest in clinical trials and has mostly been applied in cardiovascular trials. Here, we applied the win ratio approach to data from COMET, a phase 3 trial in late-onset Pompe disease, to illustrate how this approach can be used to analyze multiple endpoints in the orphan drug context. METHODS: All possible participant pairings from both arms of COMET were compared sequentially on changes at week 49 in upright forced vital capacity (FVC) % predicted and six-minute walk test (6MWT). Each participant's response for the two endpoints was first classified as a meaningful improvement, no meaningful change, or a meaningful decline using thresholds based on published minimal clinically important differences (FVC ± 4% predicted, 6MWT ± 39 m). Each comparison assessed whether the outcome with avalglucosidase alfa (AVA) was better than (win), worse than (loss), or equivalent to (tie) the outcome with alglucosidase alfa (ALG). If tied on FVC, 6MWT was compared. In this approach, the treatment effect is the ratio of wins to losses ("win ratio"), with ties excluded. RESULTS: In the 2499 possible pairings (51 receiving AVA × 49 receiving ALG), the win ratio was 2.37 (95% confidence interval [CI], 1.30-4.29, p = 0.005) when FVC was compared before 6MWT. When the order was reversed, the win ratio was 2.02 (95% CI, 1.13-3.62, p = 0.018). CONCLUSION: The win ratio approach can be used in clinical trials of rare diseases to provide meaningful insight on treatment benefits from multiple endpoints and across disease domains.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".