A Closer Look at Weak Supervision’s Limitations in WSI Recurrence Score Prediction
Bibliographic record
Abstract
Histological examination remains the gold standard for breast cancer diagnosis, prognosis assessment and treatment guidance. Commercial molecular signature test, ONCOTYPEDX®is routinely used for luminal breast cancers to predict the probabilities of response to chemotherapy and disease recurrence. We attempted to predict RS using digital pathology and Weakly Supervised (WS) attention-based models like CLAM (Clustering-constrained Attention Multiple Instance Learning) [1] and TransMIL (Transformer based Correlated Multiple Instance Learning) [2] on our in-house dataset. In tissue samples, the malignant component is haphazardly admixed with the nonmalignant component in variable proportions. This represents a challenge for the WS attention-based models to identify high-valued diagnostic/prognostic areas within whole slide images (WSIs). To address this, we propose an interactive approach with a human in the middle (supervised) by creating a user-friendly Graphical User Interface (GUI) that allows a pathologist to provide feedback to heatmaps generated by any WS attention-based model. We incorporate the feedback from the GUI as expected scores and penalize its difference with the attention scores in the successive training process (current scores). We observe an improvement in RS prediction after the pathologist’s feedback - a 5% rise in AUC and 4% in accuracy for CLAM and a 4.5% increase in AUC and 3% in accuracy for TransMIL. We analyze the generated heatmaps and notice an improvement in cosine similarity between the expected scores and attention scores before and after the feedback - 5% and 10% increase for CLAM and TransMIL, respectively. The implementation of the proposed approach and the dataset is available for download1. Our adaptive, interactive feedback system harmonizes attention scores with expert intuition and instills higher confidence in the system’s predictions. This study establishes a potent synergy between AI and pathologists.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.025 | 0.056 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.002 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".