Consensus proposal for revised International Working Group response criteria for higher risk myelodysplastic syndromes
Bibliographic record
Abstract
Myelodysplastic syndromes/myelodysplastic neoplasms (MDS) are associated with variable clinical presentations and outcomes. The initial response criteria developed by the International Working Group (IWG) in 2000 have been used in clinical practice, clinical trials, regulatory reviews, and drug labels. Although the IWG criteria were revised in 2006 and 2018 (the latter focusing on lower-risk disease), limitations persist in their application to higher-risk MDS (HR-MDS) and their ability to fully capture the clinical benefits of novel investigational drugs or serve as valid surrogates for longer-term clinical end points (eg, overall survival). Further, issues related to the ambiguity and practicality of some criteria lead to variability in interpretation and interobserver inconsistency in reporting results from the same sets of data. Thus, we convened an international panel of 36 MDS experts and used an established modified Delphi process to develop consensus recommendations for updated response criteria that would be more reflective of patient-centered and clinically relevant outcomes in HR-MDS. Among others, the IWG 2023 criteria include changes in the hemoglobin threshold for complete remission (CR), the introduction of CR with limited count recovery and CR with partial hematologic recovery as provisional response criteria, the elimination of marrow CR, and specific recommendations for the standardization of time-to-event end points and the derivation and reporting of responses. The updated criteria should lead to a better correlation between patient-centered outcomes and clinical trial results in an era of multiple emerging new agents with novel mechanisms of action.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".