Development of a Model for Predicting Early Discontinuation of Adjuvant Chemotherapy in Stage III Colon Cancer
Bibliographic record
Abstract
PURPOSE To develop a tool that can be used to predict early discontinuation of adjuvant chemotherapy among patients with stage III colon cancer. PATIENTS AND METHODS Through record linkage of Alberta administrative and tumor registry databases, we identified a cohort of individuals age ≥ 18 years who were diagnosed with stage III colon cancer and who received adjuvant chemotherapy in Alberta between 2004 and 2015. Early discontinuation was defined as receipt of < 5 months of a planned 6-month course of chemotherapy. By a systematic review of the literature and a survey of medical oncologists, the following candidate variables were identified: age (years), number of comorbidities (0, 1, ≥ 2), cancer stage (IIIC v IIIA-B), type of chemotherapy (fluorouracil, leucovorin, and oxaliplatin; capecitabine and oxaliplatin; or monotherapy), time from surgery to chemotherapy initiation (weeks), type of treatment facility (academic or community), and distance from home to treatment center (kilometers). Models developed using penalized logistic regression and the random forest algorithm were compared. Model performance was assessed using the C-statistic, Brier score, and a calibration plot. Internal validation was performed using the bootstrap method. RESULTS From an initial 3,115 patients identified, 1,378 were deemed eligible for inclusion. Of these patients, 474 patients (34.4%) failed to complete at least 5 months of chemotherapy. Although well calibrated, the penalized logistic regression model had poor discrimination (optimism-adjusted C-statistic, 0.63; 95% CI, 0.60 to 0.67). In contrast, the random forest model achieved adequate discrimination (optimism-adjusted C-statistic, 0.80; 95% CI, 0.79 to 0.82). Although the degree of calibration of the random forest was acceptable, it was slightly worse than that of the penalized logistic regression model. CONCLUSION Internal validation of our random forest model suggests that it may have clinical utility. Additional research regarding its external validation and clinical impact is needed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.011 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".