Estimating the optimal rate of adjuvant chemotherapy utilization for stage III colon cancer
Bibliographic record
Abstract
BACKGROUND: Identifying optimal chemotherapy utilization rates can drive improvements in quality of care. We report a benchmarking approach to estimate the optimal rate of adjuvant chemotherapy (ACT) for stage III colon cancer. METHODS: The Ontario Cancer Registry and linked treated records were used to identify ACT utilization. Monte Carlo simulation was used to estimate the proportion of ACT rate variation that could be due to chance alone. The criterion-based benchmarking approach was used to explore whether socioeconomic or system-level factors were associated with ACT. We also used the "pared-mean" approach to identify a benchmark population of hospitals with the highest ACT rates. RESULTS: The study population included 2801 patients; ACT was delivered to 66% (1861/2801). Monte Carlo simulation suggested that the observed component of variation (15.6%) in ACT rates was within the 95% CI (11.5%-17.3%) of what could be expected due to chance alone; the nonrandom component of ACT rate variation across hospitals was only 1.5%. There was no difference in hospital ACT rate by teaching status (P = .107), cancer center status (P = .362), or having medical oncology on site (P = .840). Unadjusted ACT rates varied across hospitals (range 44%-91%, P = .017). The unadjusted benchmark ACT rate was 81% (95%CI 76%-86%); utilization rate in non-benchmark hospitals was 65% (95%CI 63%-66%). However, after adjusting for case mix, the difference in ACT utilization between benchmark and non-benchmark populations was significantly smaller. CONCLUSIONS: We did not find any system-level factors associated with the utilization of ACT. Our results suggest that the observed variation in hospital ACT rate is not significantly different from variation due to chance alone. Using the "pared-mean" approach may significantly overestimate optimal treatment rates if case mix is not considered.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".