MétaCan
Menu
Back to cohort
Record W4205601010 · doi:10.1002/path.5864

Computer‐extracted features of nuclear morphology in hematoxylin and eosin images distinguish <scp>s</scp>tage <scp>II</scp> and <scp>IV</scp> colon tumors

2022· article· en· W4205601010 on OpenAlexaff
Neeraj Kumar, Ruchika Verma, Chuheng Chen, Cheng Lu, Pingfu Fu, Joseph Willis, Anant Madabhushi

Bibliographic record

VenueThe Journal of Pathology · 2022
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversity of Alberta
FundersNational Center for Advancing Translational SciencesNational Center for Research ResourcesNational Institute of Biomedical Imaging and BioengineeringNational Cancer InstituteNational Heart, Lung, and Blood InstituteU.S. Department of Veterans Affairs
KeywordsColorectal cancerH&E stainArtificial intelligenceStage (stratigraphy)Digital pathologyMedicineRandom forestConvolutional neural networkCancerPathologyOncologyInternal medicineComputer scienceImmunohistochemistryBiology

Abstract

fetched live from OpenAlex

We assessed the utility of quantitative features of colon cancer nuclei, extracted from digitized hematoxylin and eosin-stained whole slide images (WSIs), to distinguish between stage II and stage IV colon cancers. Our discovery cohort comprised 100 stage II and stage IV colon cancer cases sourced from the University Hospitals Cleveland Medical Center (UHCMC). We performed initial (independent) model validation on 51 (143) stage II and 79 (54) stage IV colon cancer cases from UHCMC (The Cancer Genome Atlas's Colon Adenocarcinoma, TCGA-COAD, cohort). Our approach comprised the following steps: (1) a fully convolutional deep neural network with VGG-18 architecture was trained to locate cancer on WSIs; (2) another deep-learning model based on Mask-RCNN with Resnet-50 architecture was used to segment all nuclei from within the identified cancer region; (3) a total of 26 641 quantitative morphometric features pertaining to nuclear shape, size, and texture were extracted from within and outside tumor nuclei; (4) a random forest classifier was trained to distinguish between stage II and stage IV colon cancers using the five most discriminatory features selected by the Wilcoxon rank-sum test. Our trained classifier using these top five features yielded an AUC of 0.81 and 0.78, respectively, on the held-out cases in the UHCMC and TCGA validation sets. For 197 TCGA-COAD cases, the Cox proportional hazards model yielded a hazard ratio of 2.20 (95% CI 1.24-3.88) with a concordance index of 0.71, using only the top five features for risk stratification of overall survival. The Kaplan-Meier estimate also showed statistically significant separation between the low-risk and high-risk patients, with a log-rank P value of 0.0097. Finally, unsupervised clustering of the top five features revealed that stage IV colon cancers with peritoneal spread were morphologically more similar to stage II colon cancers with no long-term metastases than to stage IV colon cancers with hematogenous spread. © 2022 The Authors. The Journal of Pathology published by John Wiley & Sons Ltd on behalf of The Pathological Society of Great Britain and Ireland.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.006
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesResearch integrity
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.686
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.006
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0000.001
Research integrity0.0000.002
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.007
GPT teacher head0.255
Teacher spread0.248 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designObservational
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations18
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueThe Journal of PathologySame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207