Comparison of Four Clinical Prognostic Scores in Patients with Advanced Gastric and Esophageal Cancer
Bibliographic record
Abstract
BACKGROUND: Prognostic scores that can identify patients at risk for early death are needed to aid treatment decision-making and patient selection for clinical trials. We compared the accuracy of four scores to predict early death (within 90 days) and overall survival (OS) in patients with metastatic gastric and esophageal (GE) cancer. METHODS: Advanced GE cancer patients receiving first-line systemic therapy were included. Prognostic risks were calculated using: Royal Marsden Hospital (RMH), MD Anderson Cancer Centre (MDACC), Gustave Roussy Immune (GRIm-Score), and MD Anderson Immune Checkpoint Inhibitor (MDA-ICI) scores. Overall survival (OS) was estimated using the Kaplan-Meier method. Cox proportional hazards models were used to analyze associations between prognostic scores and OS. The predictive discrimination was estimated using Harrell's c-index. Predictive ability for early death was measured using time-dependent AUCs. RESULTS: In total, 451 patients with metastatic GE cancer were included. High risk patients had shorter OS for all scores (RMH high- vs. low-risk median OS 7.9 vs. 12.2 months, P < .001; MDACC 6.8 vs. 11.9 months P < .001; GRIm-Score 5.3 vs. 13 months, P < .001; MDA-ICI 8.2 vs. 12.2 months, P < .001). On multivariable analysis, each prognostic score was significantly associated with OS. The GRIm-Score had the highest predictive discrimination and predictive ability for early death. CONCLUSIONS: The GRIm-Score had the highest accuracy in predicting early death and OS. Clinicians may use this score to identify patients at higher risk of early death to guide treatment decisions including clinical trial enrolment. This score could also be used as a stratification factor in future clinical trial designs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".