IMG-89. Volumetric (3D) compared with traditional bidirectional (2D) assessment in patients with Grade 2, isocitrate dehydrogenase 1/2 mutant (mIDH1/2) glioma receiving vorasidenib or placebo in the Phase 3 INDIGO trial
Bibliographic record
Abstract
Abstract While Response Assessment in Neuro-oncology (RANO) criteria remain the standard for brain tumor evaluation, 3D volumetric measurements may provide more accurate tumor assessments. We compared 2D RANO per blinded independent review committee (BIRC) with 3D volumetric RANO to evaluate correlation, progression detection and clinical outcomes in patients with mIDH1/2 diffuse glioma from INDIGO (NCT04164901). Data comparing vorasidenib (n=168) with placebo (n=163) were analyzed. Per RANO criteria for low-grade gliomas, analyses of 2D versus 3D volumetric progression-free survival (PFS) by BIRC, correlation between measurements, time to next intervention (TTNI), and response concordance were performed. Mixed models evaluated longitudinal correlation; weighted kappa statistics assessed agreement in response categorization. Strong positive correlation was observed between 2D and 3D measurements (Pearson, r=0.858; mixed model, r=0.861). 3D volumetric assessment demonstrated a stronger treatment effect (HR 0.24; 95% CI 0.15, 0.37) than 2D assessment (HR 0.35; 95% CI 0.25, 0.49). 3D volumetric assessment detected fewer progression events (vorasidenib, 16.1%; placebo, 44.2%) than 2D assessment (vorasidenib, 32.1%; placebo, 63.8%). Importantly, 3D volumetric PFS visually matched TTNI with similar treatment effect sizes. Weighted kappa analysis showed fair agreement between methods for best overall response (0.382) and all timepoint assessments (0.325). However, across all timepoint responses, 2D assessments identified more minor response (MR) and progressive disease (PD) classifications (MR: 2D, n=80; 3D n=54; PD: 2D, n=389; 3D, n=266) while 3D volumetric assessments identified more stable disease classifications (2D, n=1410; 3D, n=1552). 3D volumetric assessments demonstrated stronger treatment effects than 2D assessments, with similar treatment effect sizes for 3D PFS and TTNI. 3D volumetric assessments yielded fewer PD but more SD classifications versus 2D, suggesting a more conservative progression threshold or stable tumor burden measurement over time. Volumetric assessment may provide a more clinically meaningful determination of progression in Grade 2 glioma trials, potentially impacting future trial design and endpoint selection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".