Interpretation and Application of the International Myeloma Working Group (IMWG) Criteria: Proposal for Uniform Assessment and Reporting in Clinical Trials Based on the First Study Independent Response Adjudication Committee (IRAC) Experience
Bibliographic record
Abstract
Abstract Background: In multiple myeloma (MM), reproducible criteria of disease response and progression are critical to ensuring consistency in trial analysis and reporting. Regulatory Agencies responsible for drug approval often require clinical trials use objective endpoints that are evaluated by Independent Response Adjudication Committees (IRACs). The International Myeloma Working Group (IMWG) has developed objective criteria to define disease evaluability, response, and progression (Durie, Leukemia 2006). However, there are scenarios were IMWG criteria are ambiguous, potentially leading to inconsistency amongst IRAC members or between different IRACs when interpreting response data. To address these practical issues, we developed rules for applying IMWG response criteria to the FIRST trial, the largest study in newly-diagnosed MM (Facon, Blood 2013). Patients and Methods: FIRST is a pivotal phase III trial for previously untreated patients with MM not eligible for ASCT that enrolled 1623 patients; the primary endpoint was progression-free survival (PFS). At 12 in-person meetings between 2010-2013, the IRAC assessed eligibility, evaluability and response status of all patients after each cycle until PD or study discontinuation. These evaluations were used in the trial’s primary analysis. Response was based on central laboratory values and assessed using IMWG criteria. For circumstances where IMWG criteria were ambiguous, rules were developed through unanimous consensus of IRAC members and then applied uniformly throughout the study. Results: Rules addressing identified issues on evaluability, response and progressive disease are shown in tables 1-3. Common situations posing a need for rules concerned to measurability, missing laboratory values, timing of BM exam to assess CR, discrepancies between screening and baseline lab values or measurements in the size of extramedullary plasmacytomas Conclusions: These recommendations provide explicit descriptions of response assessment of the FIRST trial, can be used for a more uniform evaluation and reporting in future clinical studies and can assist investigators’ adherence to clinical trial requirements. Table 1. Rules for Use of Data for Evaluation Issue Recommendation Light chain (Bence-Jones) myeloma with “non-measurable” serum light chain Use only 24 hour urine M-spike value for response evaluation, except for complete response (CR) IgG, IgA or IgD myeloma with “non-measurable” serum M-spike values and measurable urine M-spike Use only urine values for response evaluation except for CR or PD Disease with “measurable” values at screening but “non-measurable” at baseline (cycle 1, day 1) All assessments not meeting CR or PD should be “non-evaluable (NE)” Missing data for 2 or more consecutive cycles Consider “NE” for the specific missing cycle assessments M-spike reported as “too small to quantitate” in responding patient Assign value of 0 to allow subsequent calculation of absolute increase to determine PD Plasmacytoma given prior radiation therapy or located only in bone Not used for response assessment, except for potential PD Table 2. Rules for Response Assessment Issue Recommendation Absence of 2 consecutive negative IFE and simultaneous <5% BMPCs CR not assigned, assess as VGPR Extramedullary plasmacytomas (EMPs) - Visits until first EMP assessment Assess as NE - Two consecutive missing EMP assessments Assess as NE - EMPs not assessed as per protocol Assess as NE (consider a sensitivity analysis (ignoring EMPs)) - Patients in serologic VGPR, with ³ 50% decrease in EMP, but still present Assess as PR, until EMPs have disappeared Table 3. Rules for Determining Progressive Disease Issue Recommendation Increase in a previously existing EMP or bone lesion as only source of PD Request verification of radiologist reports before PD is assigned Initiation of a new antimyeloma therapy before documented PD Censor at the time of last assessment before starting the new therapy PD only based on the BMPCs Determine reason for BM exam (anemia? bone pain?) before assigning PD Radiation therapy not for pre-planned reasons Assess as PD PD based on M-protein measurements with no confirmation Censor unless that PD is considered unequivocal by unanimous agreement of IRAC Disclosures Blade: Janssen: Honoraria, Research Funding; Celgene: Honoraria, Research Funding. Knop:Celgene: Honoraria. Cohen:Celgene: Honoraria. Shah:Onyx Pharmaceuticals: Consultancy, Research Funding; Celgene: Consultancy, Research Funding; Millennium Pharmaceuticals: Consultancy, Research Funding; Novartis: Consultancy, Research Funding; Array: Consultancy, Research Funding. Meyer:Celgene: Honoraria.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.896 | 0.835 |
| Meta-epidemiology (narrow) | 0.006 | 0.008 |
| Meta-epidemiology (broad) | 0.014 | 0.019 |
| Bibliometrics | 0.023 | 0.018 |
| Science and technology studies | 0.009 | 0.020 |
| Scholarly communication | 0.035 | 0.015 |
| Open science | 0.025 | 0.018 |
| Research integrity | 0.028 | 0.047 |
| Insufficient payload (model declined to judge) | 0.002 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".