MétaCan
Menu
Back to cohort

Prostate cancer risk in African American men evaluated via digital histopathology multi-modal deep learning models developed on NRG Oncology phase III clinical trials.

2022· article· en· W4281742561 on OpenAlexaff
Mack Roach, Jingbin Zhang, Andre Esteva, Osama Mohamad, Douwe van der Wal, Jeffry Simko, Sandy DeVries, Huei–Chung Huang, Edward M. Schaeffer, Todd M. Morgan, Jedidiah M. Monson, Farah Naz, James A. Wallace, Michelle Ferguson, Jean-Paul Bahary, Howard M. Sandler, Daniel E. Spratt, Stephanie L. Pugh, Phuoc T. Tran, Felix Y. Feng

Bibliographic record

VenueJournal of Clinical Oncology · 2022
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsUniversité de MontréalUniversity of SaskatchewanHorizon Health NetworkSaint John Regional Hospital
FundersNational Institutes of Health
KeywordsMedicineProstate cancerInternal medicineProportional hazards modelOncologyClinical endpointCumulative incidenceClinical trialHazard ratioCancerConfidence intervalCohort

Abstract

fetched live from OpenAlex

108 Background: Artificial intelligence (AI) tools can display racial bias as a result of existing systemic health inequities and biased datasets. We have previously developed multi-modal AI (MMAI) prognostic models based on digital pathology images from five phase III randomized radiotherapy prostate cancer trials that outperform NCCN risk groups for prediction of distant metastasis (DM), biochemical failure (BF), prostate cancer-specific mortality (PCSM) and all-cause mortality (OS). In this study, we assessed the algorithmic fairness of the locked MMAI models between African American (AA) and non-AA populations in the five randomized trials. Methods: Patients enrolled in NRG/RTOG 9202, 9408, 9413, 9910, and 0126 with digitized biopsy histopathology slides were included in this study. The locked MMAI models were applied, and subgroup analyses were conducted by comparing distributions of clinical variables and MMAI scores (medians for continuous variables and proportions for categorical variables reported), and evaluating MMAI models’ prognostic ability among AA and non-AA men. The performance of the models were compared using DM as the primary endpoint and secondary endpoints of BF, PCSM, OS (death without an event as a competing risk) with Fine-Gray or Cox Proportional Hazards models. Either Kaplan Meier or cumulative incidence estimates were computed and compared using log-rank or Gray’s test. Results: This study included 5,624 men: 932 (17%) AA, 4503 (80%) white, and 189 (3%) other races. AA had younger median age (69 vs 71 year [yr]), higher median baseline PSA (12 vs 10 ng/mL), more T1-T2a (62% vs 57%), more Gleason < 7 (42% vs 36%) and 8-10 (15% vs 12%), and more NCCN low and high risk (12% vs 10% and 41% vs 33%). AA and non-AA had estimated 5-yr BF rates 27% and 27%, 5-yr DM rates 5% and 5%, 10-yr PCSM 5% and 7%, and 10-yr OS 58% and 60%, respectively. The median (interquartile range) score of the model optimizing for 5-yr DM (5-yr DM MMAI) was 0.044 (0.037–0.059) in AA and 0.043 (0.036–0.057) in non-AA. Similarly, all other MMAI models had differences in the medians between AA and non-AA ranging from 0.001 to 0.02. For all endpoints, the 5-yr DM MMAI model showed strong prognostic signal (hazard ratio [HR] per one standard deviation increase: 1.6 for DM, 1.4 for BF, 1.6 for PCSM and 1.3 for OS, all p-values < 0.001) and had comparable trends within AA vs. non-AA in the entire cohort (e.g., HR for DM 1.4 vs 1.6). Similar results were observed for the MMAI model optimizing for 10-yr PCSM. Conclusions: To our knowledge, this represents the first comparative analyses of a digital pathology AI prognostic model in AA vs. non-AA prostate cancer patients. The prognostic performance of the AI models was found to be comparable between subgroups. Our data supports the use of these models across racial groups, though further validation in AA cohorts is ongoing.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.055
metaresearch head score (Gemma)0.051
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Meta-epidemiology (narrow), Research integrity
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Other design · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.855
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0550.051
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0060.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0010.000
Research integrity0.0000.011
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.213
GPT teacher head0.551
Teacher spread0.339 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designOther design
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations6
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Clinical OncologySame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207