Comparison of four clinical prognostic scores in patients with advanced gastric and esophageal cancer.
Bibliographic record
Abstract
4057 Background: While several clinical scoring systems exist to aid prognostication and patient (pt) selection for clinical trials in oncology, none are standardly used. We compared the ability of four prognostic scores to predict overall survival (OS) in pts with advanced gastric and esophageal (GE) cancer. Methods: Pts with advanced (unresectable or metastatic) GE cancer receiving first-line palliative-intent systemic therapy at the Princess Margaret Cancer Centre from 2007 to 2020 were included. High prognostic risk pts were identified using four scoring systems: Royal Marsden Hospital (RMH), MD Anderson Cancer Centre (MDACC), Gustave Roussy Immune Score (GRIm-S) and MD Anderson Immune Checkpoint Inhibitor (MDA-ICI) score. OS was estimated using the Kaplan-Meier method and compared between risk groups (high vs. not-high) for each scoring system using the log-rank test. Cox proportional hazards models were used to analyze the association between each prognostic score and OS, adjusting for baseline clinical factors. Harrell’s c-index was used to evaluate predictive discrimination of the models. Time-dependent AUCs were used to measure predictive ability for early death (within 90 days). Results: In total, 451 pts with advanced GE cancer were included. The median age was 59 years, 68% were male, 51% had ECOG status 0-1, 63% presented with de novo metastatic disease. The proportion of pts categorized as high risk was: RMH 25% (N=113), MDACC 13% (N=95), GRIm-S 24% (N=109), MDA-ICI 26% (N=117). In all scoring systems, high risk pts had significantly shorter OS (median OS 7.9 versus 12.2 months for RMH high vs. low risk, p<0.001; 6.8 vs. 11.9 months p<0.001 for MDACC; 5.3 vs. 13 months p<0.001 for GRIm-S; 8.2 vs. 12.2 months p<0.001 for MDA-ICI). On multivariable analysis, each prognostic score was significantly associated with OS (Table). The GRIm-S had the highest predictive discrimination (c-index 0.645 [0.612-0.678]) and highest predictive ability for early death (AUC 0.754 [0.675-0.832]). Conclusions: All four prognostic scoring systems compared had reasonable accuracy in predicting OS for patients with advanced GE cancer. The higher accuracy for predicting early death may render the GRIm-S as preferable. These tools can aid oncologists in discussions about prognosis, therapeutic decision-making and patient selection for clinical trials.[Table: see text]
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".