Data‐driven optimization of version 9 American Joint Committee on Cancer staging system for anal cancer
Bibliographic record
Abstract
INTRODUCTION: The American Joint Committee on Cancer (AJCC) staging system undergoes periodic revisions to maintain contemporary survival outcomes related to stage. Recently, the AJCC has developed a novel, systematic approach incorporating survival data to refine stage groupings. The objective of this study was to demonstrate data-driven optimization of the version 9 AJCC staging system for anal cancer assessed through a defined validation approach. METHODS: The National Cancer Database was queried for patients diagnosed with anal cancer in 2012 through 2017. Kaplan-Meier methods analyzed 5-year survival by individual clinical T category, N category, M category, and overall stage. Cox proportional hazards models validated overall survival of the revised TNM stage groupings. RESULTS: Overall, 24,328 cases of anal cancer were included. Evaluation of the 8th edition AJCC stage groups demonstrated a lack of hierarchical prognostic order. Survival at 5 years for stage I was 84.4%, 77.4% for stage IIA, and 63.7% for stage IIB; however, stage IIIA disease demonstrated a 73.0% survival, followed by 58.4% for stage IIIB, 59.9% for stage IIIC, and 22.5% for stage IV (p <.001). Thus, stage IIB was redefined as T1-2N1M0, whereas Stage IIIA was redefined as T3N0-1M0. Reevaluation of 5-year survival based on data-informed stage groupings now demonstrates hierarchical prognostic order and validated via Cox proportional hazards models. CONCLUSION: The 8th edition AJCC survival data demonstrated a lack of hierarchical prognostic order and informed revised stage groupings in the version 9 AJCC staging system for anal cancer. Thus, a validated data-driven optimization approach can be implemented for staging revisions across all disease sites moving forward.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".