Long-term trend of quality-adjusted time without symptoms or toxicities (Q-TWiST) of nivolumab+ipilimumab (N+I) versus sunitinib (SUN) for the first-line treatment of advanced renal cell carcinoma (aRCC).
Bibliographic record
Abstract
6568 Background: After a minimum follow-up of 48 months (mos), the CheckMate 214 trial (phase 3, NCT02231749) continued to demonstrate a significant overall (OS) and progression-free (PFS) survival benefit for N+I vs. SUN in aRCC patients (pts) with intermediate (I) or poor (P) International Metastatic RCC Database Consortium (IMDC) risk factors (median OS: 48.1 vs. 26.6 mos, HR: 0.65, 95% confidence interval [95% CI]: 0.54, 0.78; 48-mos PFS: 32.7% vs. 12.3%, HR: 0.74, 95% CI: 0.62, 0.88) (Albiges et al. ESMO Open 2020). To further understand the clinical benefits and risks of N+I vs. SUN, we evaluated the Q-TWiST over time using up to 57 mos of follow-up in CheckMate 214. Methods: OS was partitioned into 3 states: time with any grade 3 or 4 adverse events (TOX), time without symptoms of disease or toxicity (TWiST), and time after progression (REL). The Q-TWiST is a metric that combines the quantity and quality (i.e., “utility”) of time spent in each of the 3 states TWiST, TOX, and REL. Prior research (Revicki et al, Qual Life Res, 2006) has established that relative gains in Q-TWiST (i.e., Q-TWiST gain divided by OS in SUN) of ≥ 10% and ≥ 15% can be considered as “clinically important” and “clearly clinically important”, respectively. Non-parametric bootstrapping was used to generate 95% CIs. To observe changes in quality-adjusted survival gains over time, absolute and relative Q-TWiST were calculated up to 57 mos at intervals of 12-mos. Results: With 57-mos follow-up, compared to SUN pts, N+I pts (N = 847) had significantly longer time in TWiST state (+7.1 mos [95% CI: 4.2, 10.4]). The between-group differences in TOX state (0.3 mos [95% CI: -0.2, 0.8]) and REL state (-1.2 mos [95% CI: -4.1, 1.5]) were not statistically significant. The Q-TWiST gain in the N+I vs. SUN arms was 6.6 mos (95% CI: 4.1, 9.4), resulting in a 21.2% relative gain. Q-TWiST gains progressively increased over the follow-up period and exceeded the “clinically important” threshold around 27 mos (Table). These gains were driven by steady increases in TWiST gains from 0.4 mos (after 12 mos) to 7.1 mos (after 57 mos). Conclusions: In CheckMate 214, N+I resulted in a statistically significant and “clearly clinically important (≥ 15%)” longer quality-adjusted survival vs. SUN, which increased over the longer follow-up time. Q-TWiST gains were primarily driven by time in “good” health (i.e., TWiST), which largely resulted from the long-term PFS benefits seen for N+I vs. SUN. Clinical trial information: NCT02231749. [Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.003 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".