Discordance in Total Mesorectal Excision Specimen Grading in a Prospective Phase 2 Multicenter Rectal Cancer Trial
Bibliographic record
Abstract
OBJECTIVES: To report the results of a rigorous quality control (QC) process in the grading of total mesorectal excision (TME) specimens during a multicenter prospective phase 2 trial of transanal TME. BACKGROUND: Grading of TME specimens is based on the macroscopic assessment of the mesorectum and standardized through synoptic pathology reporting. TME grade is a strong predictor of outcomes with incomplete (IC) TME associated with increased rates of local recurrence relative to complete or near complete (NC) TME. Although TME grade serves as an endpoint in most rectal cancer trials, in protocols incorporating centralized review of TME specimens for quality assurance, discordance in grading and the management thereof has not been previously described. METHODS: A phase 2 prospective transanal TME trial was conducted from 2017 to 2022 across 11 North American centers with TME quality as the primary study endpoint. QC measures included (1) training of site pathologists in TME protocols, (2) blinded grading of de-identified TME specimen photographs by central pathologists, and (3) reconciliation of major discordance before trial reporting. Cohen Kappa statistic was used to assess agreement in grading. RESULTS: Overall agreement in grading of 100 TME specimens between site and central reviewer was rated as fair, (κ = 0.35; 95% CI: 0.10-0.61; P < 0.0001). Concordance was noted in 54%, with minor and major discordance in 32% and 14% of cases, respectively. Upon reconciliation, 13/14 (93%) major discordances were resolved. Pre versus postreconciliation rates of complete or NC and IC TME are 77%/16% and 7% versus 69%/21% and 10%. Reconciliation resulted in a major upgrade (IC-NC; N = 1) or major downgrade (NC/C-IC, N = 4) in 5 cases overall (5%). CONCLUSIONS: A 14% rate of major discordance was observed in TME grading between the site and central reviewers. The resolution resulted in a major change in final TME grade in 5% of cases, which suggests that reported rates or TME completeness are likely overestimated in trials. QC through a central review of TME photographs and reconciliation of major discordances is strongly recommended.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.101 | 0.104 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".