MétaCan
Menu
Back to cohort

A Closer Look at Weak Supervision’s Limitations in WSI Recurrence Score Prediction

2023· article· en· W4390992147 on OpenAlexaff
N. Guruprasad, Amir Akbarnejad, Gilbert Bigras, Penny J. Barnes, Nilanjan Ray

Bibliographic record

Venuenot available
Typearticle
Languageen
FieldComputer Science
TopicAI in cancer detection
Canadian institutionsDalhousie UniversityUniversity of Alberta
Fundersnot available
KeywordsComputer scienceNoticeCluster analysisArtificial intelligenceGraphical user interfaceMachine learningF1 scoreCalculatorPattern recognition (psychology)

Abstract

fetched live from OpenAlex

Histological examination remains the gold standard for breast cancer diagnosis, prognosis assessment and treatment guidance. Commercial molecular signature test, ONCOTYPEDX®is routinely used for luminal breast cancers to predict the probabilities of response to chemotherapy and disease recurrence. We attempted to predict RS using digital pathology and Weakly Supervised (WS) attention-based models like CLAM (Clustering-constrained Attention Multiple Instance Learning) [1] and TransMIL (Transformer based Correlated Multiple Instance Learning) [2] on our in-house dataset. In tissue samples, the malignant component is haphazardly admixed with the nonmalignant component in variable proportions. This represents a challenge for the WS attention-based models to identify high-valued diagnostic/prognostic areas within whole slide images (WSIs). To address this, we propose an interactive approach with a human in the middle (supervised) by creating a user-friendly Graphical User Interface (GUI) that allows a pathologist to provide feedback to heatmaps generated by any WS attention-based model. We incorporate the feedback from the GUI as expected scores and penalize its difference with the attention scores in the successive training process (current scores). We observe an improvement in RS prediction after the pathologist’s feedback - a 5% rise in AUC and 4% in accuracy for CLAM and a 4.5% increase in AUC and 3% in accuracy for TransMIL. We analyze the generated heatmaps and notice an improvement in cosine similarity between the expected scores and attention scores before and after the feedback - 5% and 10% increase for CLAM and TransMIL, respectively. The implementation of the proposed approach and the dataset is available for download1. Our adaptive, interactive feedback system harmonizes attention scores with expert intuition and instills higher confidence in the system’s predictions. This study establishes a potent synergy between AI and pathologists.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.025
metaresearch head score (Gemma)0.056
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.025
Threshold uncertainty score0.130

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0250.056
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0010.001
Science and technology studies0.0010.001
Scholarly communication0.0020.003
Open science0.0030.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0010.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.082
GPT teacher head0.271
Teacher spread0.188 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same topicAI in cancer detectionFrench-language works237,207