Interpretable Multiple Instance Learning for Hematologic Diagnosis from Peripheral Blood Smears
Bibliographic record
Abstract
Accurate diagnosis of hematologic malignancies from peripheral blood smears (PBSs) requires integrating cellular morphology and composition across hundreds of white blood cells. Existing approaches primarily automate single-cell classification and do not provide whole-slide diagnostic predictions. We present a full network that utilizes a highly performative cell-based encoder (DeepHeme) for feature extraction paired with our weakly supervised framework using attention-based multiple instance learning (MIL) that we call CAREMIL (Cell AggRegation, Explainable, Multiple Instance Learning). Upon evaluating various popular image encoders and MIL architectures, the combination of DeepHeme and CAREMIL is the best performing pipeline on our disease classification task. CAREMIL proves to be a robust aggregation function that outperforms the most commonly used slide level aggregation function (gated multiple instance learning) across several encoder types. The greatest improvements in performance gain with CAREMIL is observed when using out-of-domain encoders, including an encoder trained on ImageNet and leading open-source pathology foundational models (UNI2 and Virchow2). CAREMIL plus DeepHeme achieves the highest diagnostic performance across acute leukemia (AML), myelodysplastic syndromes (MDS), and hairy cell leukemia (HCL) (AUROCs 0.999, 0.891, and 0.945, respectively), and identifies AML disease even in cases with minimal or absent circulating blasts. Attention values assigned by CAREMIL highlight diagnostically relevant cells and reveal disease-specific morphometric signatures, enabling biological interpretability and case-level insight. CAREMIL remains robust to misclassified cell types by the cell image encoder and does not require explicit cell-level supervision. These findings position CAREMIL as an effective and interpretable multiple instance learning framework for hematologic slide diagnosis, with potential to extend to bone marrow aspirates, cytology, and other liquid biopsy specimens, and to support a broader shift toward quantitative, morphology-informed diagnostics in hematology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.007 | 0.006 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".