MétaCan
Menu
Back to cohort

Abstract P4-02-01: Analytical validation of an automated digital scoring protocol for Ki67: International multicenter collaboration study

2019· article· en· W2913972996 on OpenAlexaff
Balázs Ács, SC Leung, Vasiliki Pelekanou, Yibing Bai, Sandra Martínez-Morilla, Maria Toki, MC Chang, Abhi Gholap, A Jadhav, JC Hugh, Gilbert Bigras, Arvydas Laurinavičius, Renaldas Augulis, Richard M. Levenson, A Todd, Tammy Piper, Shakeel Virk, Bert van der Vegt, DF Hayes, Mitch Dowsett, TO Nielsen, DL Rimm

Bibliographic record

VenueCancer Research · 2019
Typearticle
Languageen
FieldMedicine
TopicRadiomics and Machine Learning in Medical Imaging
Canadian institutionsSinai Health SystemUniversity of TorontoUniversity of British Columbia
Fundersnot available
KeywordsReproducibilityBreast cancerProtocol (science)MedicineStandardizationMedical physicsTissue microarrayDigital image analysisCancerPathologyArtificial intelligenceComputer scienceInternal medicineStatisticsMathematicsComputer vision

Abstract

fetched live from OpenAlex

Abstract Background/Goal: Ki67 expression has been a valuable prognostic marker in breast cancer, but has not seen broad adoption due to lack of standardization between institutions. Automation could represent a solution. Here we tested 3 automated digital image analysis (DIA) platforms including an open source platform to: (i) Investigate the reproducibility of Ki67 measurement across platforms with supervised classifiers performed by the same operator and by multiple operators. (ii) Compare accuracy of the 3 DIA platforms against outcome (prognostic potential). (iii) Assess inter-laboratory reproducibility of a calibrated DIA tool to evaluate Ki67 in breast cancer among 10 participating labs of the International Ki67 in Breast Cancer Working Group (IKWG). Methods: The Mib-1 antibody (Dako) was used to detect Ki67 (dilution 1:100). HALO (H) (IndicaLabs), QuantCenter (QC) (3DHistech), QuPath (QP) (open-source software) digital image analysis (DIA) platforms were used to evaluate Ki67 expression. As a ground truth, we evaluated Ki67 LI with meticulous manual tissue segmentation using the Spectrum Webscope (SW) (Aperio). Calibration was performed using 30 ER+ breast cancer cases from phase 3 of the IKWG initiative where blocks were centrally cut and stained for Ki67. The inter-laboratory analysis was done with 10 participating laboratories divided into 2 groups where members within the same group were given the same set of images. The outcome cohort consisted of 149 breast cancer cases from the Yale Pathology archives in tissue microarray format. Intra-class correlation coefficient (ICC) was used to measure reproducibility with the pre-specified criterion for success being to exceed 0.80. Kaplan-Meier analysis supported with log-rank test was performed to assess prognostic potential. Results: All 3 DIA platforms showed excellent inter-platform reproducibility (ICC: 0.933, CI: 0.879-0.966). Also, excellent reproducibility was found between all DIA platforms and the reference standard Ki67 values of SW (QP ICC: 0.970, CI: 0.936-0.986; H ICC: 0.968, CI: 0.933-0.985; QC ICC: 0.964, CI: 0.919-0.983). The intra-DIA reproducibility was also excellent for all platforms (QP ICC: 0.992, CI: 0.986-0.996; H ICC: 0.972, CI: 0.924-0.988; QC ICC: 0.978, CI: 0.932-0.991). Comparing each DIA against outcome, the hazard ratios were similar (QP=3.309, H=3.077, QC=3.731). The inter-operator reproducibility was particularly high (ICC: 0.962-0.995). As QP is open source software and also showed the lowest intra-DIA platform variability, we selected the QP platform to investigate inter-laboratory reproducibility among 10 IKWG labs. The different-section ICC across the 10 labs was 0.974 (CI: 0.954 - 0.986). The same-section ICC estimate was 0.984 (CI: 0.971-0.992) for group 1 and 0.978 (CI: 0.956-0.989) for group 2. Conclusions: Our results showed outstanding reproducibility both within and between DIA platforms. We also found the platforms essentially indistinguishable with respect to prediction of breast cancer patient outcome. Automated Ki67 evaluation using a calibrated, open-source DIA platform (QuPath) met the pre-specified criterion of success in the multi-institutional setting. Assessment of clinical utility is planned. Citation Format: Acs B, Leung SC, Pelekanou V, Bai Y, Martinez-Morilla S, Toki M, Chang MC, Gholap A, Jadhav A, Hugh JC, Bigras G, Laurinavicius A, Augulis R, Levenson R, Todd A, Piper T, Virk S, van der Vegt B, Hayes DF, Dowsett M, Nielsen TO, Rimm DL. Analytical validation of an automated digital scoring protocol for Ki67: International multicenter collaboration study [abstract]. In: Proceedings of the 2018 San Antonio Breast Cancer Symposium; 2018 Dec 4-8; San Antonio, TX. Philadelphia (PA): AACR; Cancer Res 2019;79(4 Suppl):Abstract nr P4-02-01.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.075
metaresearch head score (Gemma)0.080
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Bench or experimental · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.075
Threshold uncertainty score0.394

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0750.080
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0020.002
Science and technology studies0.0020.002
Scholarly communication0.0030.001
Open science0.0030.003
Research integrity0.0020.001
Insufficient payload (model declined to judge)0.0030.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.077
GPT teacher head0.532
Teacher spread0.455 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designBench or experimental
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2019
Admission routes1
Has abstractyes

Explore more

Same venueCancer ResearchSame topicRadiomics and Machine Learning in Medical ImagingFrench-language works237,207