Bibliographic record
Abstract
In recent years, the field of immune thrombocytopenia (ITP) has experienced a resurgence of randomized clinical trials (RCTs) heralded by the introduction of new treatments including rituximab and the thrombopoietin receptor agonists (romiplostim and eltrombopag). One tenet of RCT methodology is that outcomes should be clinically meaningful, otherwise the large investment of resources, time, and money may not yield impactful results. Clinicians and methodologists working in ITP have struggled with this basic principal as they try to foster research that is both clinically important and operationally feasible. The debate boils down to whether platelet count or bleeding should be used as the endpoint in clinical trials. In the clinic, the platelet count defines the diagnosis and is used to evaluate response to therapy. Indeed, thrombocytopenia is the only defining feature of ITP and ancillary tests are recommended to exclude other underlying causes [1]. However, what drives management decisions is the prevention or treatment of bleeding; and ultimately it is bleeding, along with death, quality of life, and avoidance of other ITP treatments, that is most important for patients. In this issue of the American Journal of Hematology, Altomare et al. [2] provide a critique of ITP trials investigating the use of romiplostim and eltrombopag. The authors argue that use of the platelet count as the primary outcome in those trials was a valid surrogate in the absence of data on bleeding or deaths—events that are so infrequent in this population that a trial powered on those endpoints would likely not be feasible. They justify their argument based on the fact that these data were sufficient for regulatory agency approvals and provided the basis for guideline recommendations [1]. For any surrogate outcome to be considered valid, it must correlate with, predict for, and fully capture the net effect of treatment of the true clinical outcome [3]. If we apply the first two of these criteria to ITP trials, then for the platelet count to be a valid surrogate, thrombocytopenia must correlate with and predict for bleeding. Although this relationship seems intuitive, there is a surprising lack of data to estimate odds ratios or relative risks of bleeding with severe degrees of thrombocytopenia in patients with ITP. These data would be important, as the strength of the association is reflected in statistics in addition to biological plausibility [4]. In other populations, including patients with chemotherapy-induced thrombocytopenia, the risk of bleeding appears to increase with platelet counts below 5–10 × 109/L [5]; and a study of “a number of thrombocytopenic patients with a variety of disorders” found that the risk of major bleeding was minimal with platelet counts over 10 × 109/L [6]. In a study of childhood ITP, almost all serious bleeds occurred at platelet counts less than 20 × 109/L, but some children with extreme thrombocytopenia never bled [7]. Extrapolating from these data, one can anticipate that in ITP, platelet counts above 30 × 109/L do not correlate with bleeding and that platelet counts below 10 × 109/L have some correlation; however, the details of this relationship have not been well characterized. On the other hand, there are several problems with the use of bleeding as an outcome in ITP trials. For one, there is currently no accepted ITP-specific bleeding measurement tool for children and adults. Although several ITP bleeding scales have been developed [7-9], none have been widely endorsed, perhaps because of the need for further validation. To capture bleeding in ITP trials, investigators have generally used the Common Terminology Criteria for Adverse Events [10] or the WHO bleeding scale [11]. The WHO scale was developed for trials in cancer patients; it uses nonspecific descriptive terms such as “mild” and “gross” bleeding without explicit definitions, and it has not been formally evaluated for validity or reliability [12]. Moreover, defining major bleeding as WHO Grade 2–4 raises concerns about the use of a nonserious event (Grade 2 bleeding) in this composite outcome measure [13]. New initiatives are currently underway to develop a comprehensive bleeding tool that can be used across ITP studies. The second problem with the use of bleeding as the outcome in ITP trials is that serious bleeding is infrequent. Investigators have estimated that a trial designed to reduce the risk of bleeding in children with ITP from 3% (baseline risk) to 1% would require 1,730 patients [14], an unlikely proposition for such a rare disease. Thus, it is not a surprise that a meta-analysis of five romiplostim and eltrombopag RCTs (n = 808 patients) did not find a significant improvement in bleeding events; however, this does not mean that the drugs were not effective [15]. So, where does that leave us? On the one hand, the platelet count may not correlate with or predict for bleeding (and it certainly does not fully capture the net effect of treatment). On the other hand, bleeding is difficult to measure, and a traditional RCT powered on serious bleeding events would be infeasible. The answer likely lies somewhere between the purist views of the clinician and the methodologist: platelet counts should continue to be used as outcomes in ITP trials, supplemented by carefully collected data on bleeding using a validated bleeding tool. Other clinically important outcomes should include quality of life and concomitant (or rescue) ITP therapy; however, criteria for administering rescue treatment or reducing concomitant ITP therapy must be carefully prescribed in the trial protocol. In addition, ITP research is well suited to other methods of comparative effectiveness research, which may include retrospective analyses of clinical or administrative databases, longitudinal registries, and “pragmatic” RCTs, which require a simple study design with minimal data collection, and the participation of numerous centers worldwide.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.004 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".