Utility of quantitative pathologic analysis of pT1 colorectal carcinomas to improve prediction of lymph node metastasis
Bibliographic record
Abstract
According to the National Comprehensive Cancer Network (NCCN), submucosally invasive (pT1) colorectal carcinomas (CRCs) should be evaluated for tumor grade, lymphatic invasion, and tumor budding to determine the risk of lymph node metastasis. The presence of any one of these high-risk features is an indication for surgery in endoscopically removed pT1 CRCs. In this study, we determined if quantitative pathologic analysis with the QuantCRC algorithm can augment NCCN risk stratification in a multi-institutional cohort of 512 surgically resected pT1 CRC. LASSO regression identified %high-grade, %inflammatory stroma (stromal area), and %tumor budding/poorly differentiated clusters (%TB/PDC) as important QuantCRC features and were used in subsequent logistic regression analysis. Five logistic regression models were built using NCCN and QuantCRC variables, with the combined NCCN + QuantCRC model providing the highest Area Under the Curve (AUC) of 0.74 (95% CI 0.68-0.81). A predicted probability cutoff of 0.092 provided a sensitivity of 78.3% and specificity of 62.1% in the NCCN + QuantCRC model with a 24.3% rate of lymph node positivity for high-risk (HR) tumors compared to 5.2% for low-risk (LR) CRCs. Fifteen pT1 CRCs were reclassified from NCCN LR to NCCN + QuantCRC HR and 3/15 (20%) demonstrated lymph node positivity. The median predicted probability of lymph node metastasis in the NCCN + QuantCRC model was used to define two HR groups (HR1: 0.092-0.218 and HR2: > 0.218). HR2 CRCs had a rate of lymph node positivity of 31.5% compared to 17.1% for HR1 CRCs (P = 0.02). Lastly, the NCCN + QuantCRC model was validated in a cohort of 29 endoscopically resected pT1 CRCs followed by surgical resection. In the NCCN + QuantCRC model, the 8 pN + CRCs in this cohort had a higher median predicted probability of lymph node metastasis compared to 21 pN0 CRCs (0.219 vs. 0.080, P = 0.04). In summary, the addition of variables from QuantCRC can improve risk stratification of pT1 CRCs over NCCN criteria alone.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".